Skip to content
DigitalTatva
Pricing and Plans

Claude API Pricing: Opus, Sonnet and Haiku Token Costs Explained

Claude API pricing for Opus 5.5, Sonnet 5.5, Haiku 4.5 and Fable 5.1, with caching and Batch discounts, extra fees and worked cost examples for real apps.

Claude API Pricing: Opus, Sonnet and Haiku Token Costs Explained
On this page
  1. Key takeaways
  2. Claude API pricing table for every current model
  3. Each Claude model explained: price, strengths and who it suits
  4. What the Claude API really costs: worked examples
  5. Prompt caching and Batch discounts in detail
  6. Extra Claude API costs to budget for
  7. Claude API vs Claude subscriptions
  8. How Claude API pricing compares with OpenAI and Gemini
  9. Which Claude model should you choose?
  10. Frequently asked questions
  11. Our verdict on Claude API pricing

Claude API pricing is billed per million tokens, and as of September 2026 the main models cost: Claude Opus 5.5 at $4 input and $20 output, Claude Sonnet 5.5 at $2 input and $10 output, and Claude Haiku 4.5 at $1 input and $5 output. The top-end Claude Fable 5.1 costs $10 input and $50 output.

Anthropic also offers some of the most generous discounts in the market: cache reads as low as 0.025x the input price, 50% off with the Batch API, and no surcharge for using the full 1 million token context window. This guide covers every rate, the caching math, the extra fees and which model to pick for your project.

Key takeaways

  • Opus 5.5 costs $4 / $20 per 1M tokens and is Anthropic’s recommended starting point for most workloads.
  • Sonnet 5.5 costs $2 / $10 and Haiku 4.5 costs $1 / $5, so Sonnet is the value pick for high-volume apps.
  • Cache reads cost just $0.20 per 1M on Opus 5.5 (a 95% discount) and a 5-minute cache pays for itself after one read.
  • The Batch API halves both input and output prices, for example Sonnet 5.5 drops to $1 / $5.
  • Claude 4.6 and later models use the full 1M context at standard rates, with no long-context surcharge.

Claude API pricing table for every current model

These are Anthropic’s standard rates in USD per 1 million tokens, read from the official Claude API pricing page. If you are new to token billing, start with our plain English guide to tokens.

Model Input Output 5-min cache write 1-hour cache write Cache hit Context
Claude Fable 5.1 $10.00 $50.00 $12.50 $20.00 $0.25 1M
Claude Opus 5.5 $4.00 $20.00 $5.00 $8.00 $0.20 1M
Claude Sonnet 5.5 $2.00 $10.00 $2.50 $4.00 $0.20 1M
Claude Haiku 4.5 $1.00 $5.00 $1.25 $2.00 $0.10 200K

Fable 5.1, Opus 5.5 and Sonnet 5.5 can each return up to 128K output tokens in one response. Haiku 4.5 tops out at 64K output tokens and a 200K context window. The model IDs you use in code are claude-fable-5-1, claude-opus-5-5, claude-sonnet-5-5 and claude-haiku-4-5-20251001.

Bar chart of Claude API output prices per million tokens: Fable 5.1 $50, Opus 5.5 $20, Sonnet 5.5 $10, Haiku 4.5 $5
Claude API output price by model

Each Claude model explained: price, strengths and who it suits

Anthropic’s own advice on its models overview page is to “start with Claude Opus 5.5 for most workloads”. That is unusual: most vendors steer you to their mid-tier model. Opus 5.5 is priced at a level that makes it practical for everyday production use, and its cache hit price of $0.20 per million is only 0.05x its input rate.

Choose Opus 5.5 for coding agents, long-form writing, document analysis and anything where you want strong reasoning without paying flagship prices. In everyday use, its long drafts tend to need fewer edits, which saves human time as well as tokens. Our full Claude review covers its quality in more depth.

Claude Sonnet 5.5: the high-volume value pick ($2 / $10)

Sonnet 5.5 costs half of Opus 5.5 on both input and output, with the same 1M context window and 128K output limit. It is the right choice when you process a lot of requests and the tasks are well defined: support replies, product descriptions, summaries, RAG answers over your knowledge base.

Note that Sonnet 5.5 and Opus 5.5 share the same $0.20 cache hit price. If most of your input is a cached prefix, the gap between the two narrows and Opus may be worth testing.

Claude Haiku 4.5: the budget option ($1 / $5)

Haiku 4.5 is Anthropic’s cheapest current model. It is fast and well suited to classification, tagging, routing, extraction and short replies. Its limits are tighter (200K context, 64K output), and at $1 input it is not as cheap as the budget models from some rivals, so compare carefully if pure cost is your goal.

Claude Fable 5.1: the premium tier ($10 / $50)

Fable 5.1 is the most expensive model on the list at $10 input and $50 output. It has the deepest cache discount (reads at 0.025x input, so $0.25 per million) and a 1M context window. Consider it only for tasks where Opus 5.5 falls short in your own tests. For most teams, Opus 5.5 delivers far better value.

Older Claude models still listed

Anthropic still lists earlier models with their original prices. Useful to know if you maintain an older integration:

  • Opus 5 and Opus 4.5 to 4.8: $5 input / $25 output (more than Opus 5.5).
  • Opus 4 and 4.1: $15 / $75.
  • Sonnet 5: $2 / $10. Sonnet 4 to 4.6: $3 / $15.
  • Haiku 3.5: $0.80 / $4. Mythos 5.1: $10 / $50.

Save money: If you are still on Opus 4.x or Sonnet 4.x, switching to Opus 5.5 or Sonnet 5.5 cuts your per-token price while moving you to a newer model. Opus 4.1 at $15 / $75 costs almost four times as much as Opus 5.5.

What the Claude API really costs: worked examples

Here is the arithmetic for common workloads. For a step-by-step budgeting method, see our guide on estimating AI API costs before you build.

Example 1: a support assistant with 1,000 chats a day

Each chat sends about 2,000 input tokens and returns 500 output tokens. Daily totals: 2M input and 0.5M output tokens.

Model Calculation Per day Per 30 days
Fable 5.1 (2 x $10) + (0.5 x $50) $45.00 $1,350
Opus 5.5 (2 x $4) + (0.5 x $20) $18.00 $540
Sonnet 5.5 (2 x $2) + (0.5 x $10) $9.00 $270
Haiku 4.5 (2 x $1) + (0.5 x $5) $4.50 $135

Example 2: the same assistant with prompt caching on Opus 5.5

Suppose 1,500 of the 2,000 input tokens are a fixed system prompt and product FAQ. With caching, those tokens are read at the $0.20 cache hit price instead of $4.

  • Cached reads: 1.5M x $0.20 = $0.30
  • Fresh input: 0.5M x $4 = $2.00
  • Output: 0.5M x $20 = $10.00
  • Total: about $12.30 per day plus occasional cache writes, versus $18 without caching.

Each cache write of that 1,500-token prefix on a 5-minute cache costs 1,500 x $5 per million, which is less than one cent. Anthropic says a 5-minute write breaks even after a single read, so any busy app comes out ahead. Our prompt caching explainer shows how to structure prompts to get these hits.

Example 3: analyzing a 600,000-token document

On Opus 5.5, one request with 600K input tokens and 8K output tokens costs (0.6 x $4) + (0.008 x $20) = $2.40 + $0.16 = $2.56. Because Claude 4.6 and later models have no long-context surcharge, there is no multiplier for crossing a size threshold. That is a real advantage over providers that double input rates on very long prompts. Learn more in our context window guide.

Example 4: an overnight batch job

You rewrite 5,000 product descriptions, each 4,000 tokens in and 400 tokens out, for 20M input and 2M output tokens. On Sonnet 5.5 at standard rates: (20 x $2) + (2 x $10) = $60. Through the Batch API at $1 / $5: (20 x $1) + (2 x $5) = $30.

Prompt caching and Batch discounts in detail

How Claude prompt caching is priced

  • 5-minute cache write: 1.25x the base input price, valid for 5 minutes. Breaks even after one read.
  • 1-hour cache write: 2x the base input price, valid for 1 hour. Breaks even after two reads.
  • Cache read: 0.1x base input on most models, 0.05x on Opus 5.5, and 0.025x on Fable 5.1 and Mythos 5.1.

Use the 5-minute cache for steady traffic like chatbots. Use the 1-hour cache when requests come in bursts with gaps, such as a team that analyzes the same large document several times over an afternoon.

How the Batch API works

The Batch API takes 50% off both input and output tokens in exchange for asynchronous processing. Batch prices are $2 / $10 for Opus 5.5, $1 / $5 for Sonnet 5.5 and $0.50 / $2.50 for Haiku 4.5. Caching discounts still apply alongside batch pricing. For a side-by-side of batch programs across vendors, read our Batch API savings guide.

Extra Claude API costs to budget for

Item Price When it applies
Web search tool $10 per 1,000 searches When Claude searches the web for you
Code execution $0.05 per hour after 50 free hours per day Running code in Anthropic’s sandbox
Managed Agents $0.08 per session-hour Hosted agent sessions
US-only inference 1.1x on all token categories Setting inference_geo to US on 4.6 and later models

Watch out: Long outputs are the main hidden cost. On every current Claude model, output tokens cost five times as much as input tokens. Ask for concise formats and set a sensible max_tokens value on each request.

Claude API vs Claude subscriptions

The API and the Claude apps are billed separately. A Claude Pro subscription ($20 a month, or $17 a month on annual billing) gives you chat and Claude Code with session-based limits that reset every five hours, but no API credits. If you want to compare the app plans, see our Claude pricing guide and our Claude Pro vs Max comparison.

There are two useful bridges between the two worlds. Claude Code can be switched to API credits through the Console for pay-as-you-go use, which suits developers who hit plan limits. And Claude Enterprise is priced at $20 per seat per month billed annually, plus usage at API rates. If you keep running into caps on a plan, our guide to Claude usage limits explains when switching to the API makes sense.

How Claude API pricing compares with OpenAI and Gemini

Model Input / 1M Output / 1M Long-context surcharge
Claude Opus 5.5 $4.00 $20.00 None
OpenAI GPT-6 Astra $10.00 $50.00 Yes, above 272K input
Claude Sonnet 5.5 $2.00 $10.00 None
OpenAI GPT-6.1 Sol $2.00 $10.00 Yes, above 272K input
Gemini 3.1 Pro (preview) $2.00 ($4.00 over 200K) $12.00 ($18.00 over 200K) Yes, above 200K

Sonnet 5.5 matches GPT-6.1 Sol on list price, while Opus 5.5 is much cheaper than GPT-6 Astra. On the budget end, Haiku 4.5 costs more than OpenAI’s GPT-6 Luna and Google’s Flash-Lite models, so Claude is rarely the cheapest choice for bulk classification. See our OpenAI API pricing guide and Gemini API pricing guide for the full lists, or our ChatGPT vs Claude comparison for quality differences.

Which Claude model should you choose?

  • Developers building coding tools or agents: Opus 5.5. It is Anthropic’s recommended default and handles complex, multi-step work well. Our list of the best AI coding assistants shows how Claude-based tools compare.
  • Startups with a chatbot or content feature: Sonnet 5.5 with a cached system prompt.
  • Bulk data processing on a budget: Haiku 4.5 through the Batch API at $0.50 / $2.50.
  • Legal, research or document-heavy teams: Opus 5.5 for its 1M context with no surcharge.
  • Writers and marketers who only need chat: skip the API and use a Claude plan. Our Claude writing prompts will get you more from it.

Pros

  • Opus-class quality at $4 / $20, far below other flagships
  • Deep cache discounts, down to 0.025x on Fable 5.1
  • No long-context surcharge on 4.6 and later models
  • Batch API halves both input and output

Cons

  • Haiku 4.5 is pricier than rival budget models
  • Haiku is limited to 200K context and 64K output
  • Web search and US-only inference add extra fees
  • Billed in USD only, with no rupee API pricing

Frequently asked questions

How much does the Claude API cost per request?

It depends on the model and token count. A typical request with 2,000 input tokens and 500 output tokens costs about $0.018 on Opus 5.5, $0.009 on Sonnet 5.5 and $0.0045 on Haiku 4.5. Caching a repeated prompt prefix and using the Batch API can reduce these costs substantially, while long outputs push them up because output is five times the input rate.

Is Claude API cheaper than OpenAI API?

At the mid tier they are equal: Claude Sonnet 5.5 and GPT-6.1 Sol both cost $2 input and $10 output per million tokens. At the top tier, Claude Opus 5.5 at $4 / $20 is much cheaper than GPT-6 Astra at $10 / $50. At the budget tier, OpenAI’s GPT-6 Luna is far cheaper than Claude Haiku 4.5.

Does Claude Pro include API credits?

No. Claude Pro and Max subscriptions cover the Claude apps and Claude Code with usage limits, but API usage is billed separately through the Claude Console at per-token rates. You can, however, switch Claude Code to API credits for pay-as-you-go use. Enterprise plans combine a seat fee with usage charged at API rates.

Is there a long-context surcharge on the Claude API?

Not on current models. Anthropic says Claude 4.6 and later models get the full 1M-token context window at standard pricing, with no long-context surcharge, and caching and batch discounts still apply. That covers Opus 5.5, Sonnet 5.5 and Fable 5.1. Haiku 4.5 has a smaller 200K context window.

How much does the Claude API cost in India?

Anthropic lists API prices in US dollars, and we found no rupee price list for API tokens. Rupee pricing launched in July 2026 applies to the Claude app plans, not the API. Indian developers pay the USD token rates by card, so the final rupee cost depends on exchange rates, bank fees and applicable taxes.

Our verdict on Claude API pricing

Claude’s 2026 pricing is strongest in the middle and at the top. Opus 5.5 gives you a flagship-class model at $4 / $20, Sonnet 5.5 matches the market rate for mid-tier models, and the lack of a long-context surcharge makes Claude a natural fit for big documents and codebases. Add caching and batching and your effective rate falls well below list price.

If raw cost per token for simple tasks is all that matters, look at the cheapest OpenAI or Gemini models instead. For everything else, start with Opus 5.5 as Anthropic suggests, then move high-volume paths to Sonnet 5.5. Prices can change, so always confirm current rates on Anthropic’s official pricing page before you finalize a budget.

Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.

Written by

Ketan Parmar

Ketan Parmar has spent more than 15 years in digital marketing, helping brands grow through SEO, Google Ads, Meta Ads, content strategy and social media. Today he focuses on AI search visibility: how businesses get found and recommended in ChatGPT, Gemini, Perplexity and Google's AI answers.

Get the weekly AI tools brief

New tools, price changes and money-saving deals. One email a week, no spam.