The current price table

Anthropic lists four recommended models and keeps older models available at their own rates. All figures below are US dollars per million tokens, taken from Anthropic's pricing documentation on 2026-10-04. Prices change; check the source before you budget against them.

The newest generation is cheaper per token than the one before it. Opus 5.5 lists at $4 input and $20 output against $5 and $25 for Opus 5, Opus 4.8, and Opus 4.5. Sonnet 5.5 and Sonnet 5 list at $2 and $10 against $3 and $15 for Sonnet 4.6 and 4.5. Opus 4.1 and Opus 4 remain at $15 and $75, which makes them the most expensive models still on the list and the first place to look in an older codebase.

  • Claude Fable 5.1: $10 input, $50 output, $0.25 cache read.
  • Claude Opus 5.5: $4 input, $20 output, $0.20 cache read.
  • Claude Sonnet 5.5: $2 input, $10 output, $0.20 cache read.
  • Claude Haiku 4.5: $1 input, $5 output, $0.10 cache read.
  • Claude Opus 5, 4.8, 4.7, 4.6, 4.5: $5 input, $25 output, $0.50 cache read.
  • Claude Sonnet 4.6, 4.5, 4: $3 input, $15 output, $0.30 cache read.
  • Claude Opus 4.1 and Opus 4: $15 input, $75 output, $1.50 cache read.

The tokenizer changes what a token is worth

Anthropic's pricing page states that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, with the exact increase depending on content. Sonnet 4.6 and earlier use the previous tokenizer. Per-token prices across that line are therefore not directly comparable.

A worked illustration: if your prompts produce 30% more tokens on Sonnet 5.5 than on Sonnet 4.6, the same text costs about $2.60 in Sonnet 5.5 input charges for every $3.00 it cost on Sonnet 4.6. That is a saving of roughly 13%, not the 33% the list prices suggest. Your own ratio could be higher or lower. The way to find it is to run a sample of real prompts through the token counting endpoint on both models before you migrate and before you update a forecast.

The same effect applies to output. A model that writes the same answer in more tokens bills more for it, so compare cost per completed task on your own traffic rather than cost per million tokens on a pricing page.

Modifiers that move the bill

Cache reads are where the newest models differ most. A cache read costs 0.1x the base input price on most models, 0.05x on Opus 5.5, and 0.025x on Fable 5.1. The result is that Opus 5.5 and Sonnet 5.5 both charge $0.20 per million cached input tokens. On a context-heavy workload, the gap between the two models is mostly in uncached input and output. Writes cost extra: 1.25x base input for a 5-minute cache and 2x for a 1-hour cache, so Anthropic notes the 5-minute cache pays off after one read and the 1-hour cache after two.

Consider a request with 20,000 input tokens, 18,000 of them read from cache, and 1,000 output tokens. On Sonnet 5.5 it costs about $0.0176, against $0.05 with no caching. On Opus 5.5 the same request costs about $0.0316 cached and $0.10 uncached. A cached Opus 5.5 request in this shape costs less than an uncached Sonnet 5.5 request, which is a reason to fix caching before downgrading models. These figures leave out the initial cache writes, which a steady workload spreads across many reads.

The other modifiers are flat multipliers or fees. Each stacks with the others unless noted.

  • Batch API: 50% off input and output on every listed model. Not available with fast mode or inside Claude Managed Agents sessions.
  • Fast mode (research preview, first-party API only): Opus 5.5 at $8 input and $40 output; Opus 5 and Opus 4.8 at $10 and $50.
  • US-only inference (inference_geo set to "us"): 1.1x on all token categories for Claude 4.6 and later models.
  • Long context: Claude 4.6 and later models bill the full 1M-token window at standard rates.
  • Web search: $10 per 1,000 searches plus tokens. Web fetch: tokens only.
  • Code execution: 1,550 free hours per organization per month, then $0.05 per hour per container; free when used alongside web search or web fetch tools.
  • Claude Managed Agents: token rates plus $0.08 per session-hour while running.

Estimating your own cost

Tool use adds tokens before your prompt does. Declaring any tool adds a system prompt of 286 tokens on Opus 5.5 and Sonnet 5.5, 496 on Haiku 4.5 with auto tool choice, and 675 on Opus 4.7, plus the tool names, descriptions, and schemas. Anthropic's computer use toolset adds about 4,500 input tokens per request and the browser use toolset about 6,600. An agent that resends those definitions on every step pays for them on every step unless they sit in a cached prefix.

A reliable estimate needs four numbers per workload: input tokens on the target model's tokenizer, the share of input that is a cached prefix, output tokens, and request volume. Multiply each token class by its rate, apply batch or residency multipliers where they apply, and add server-tool fees. Then compare against what the invoice says, because retries, failed calls, and evaluator traffic all bill at the same rates and rarely appear in a spreadsheet model.

Pricing pages describe rates, not spend. To know which feature, team, or customer drives the bill, every request needs an owner tag and a recorded cost at the time it runs. That can live in your own request ledger, an LLM gateway such as LiteLLM, or a FinOps layer such as FrugalAI. Partner clouds are separate: Amazon Bedrock and Google Cloud publish their own Claude pricing and invoice you directly.

Frequently asked questions

How much does the Claude API cost per million tokens?

As of 2026-10-04, Anthropic lists Claude Fable 5.1 at $10 input and $50 output, Claude Opus 5.5 at $4 and $20, Claude Sonnet 5.5 at $2 and $10, and Claude Haiku 4.5 at $1 and $5 per million tokens. The Batch API cuts both rates by 50%.

Is Sonnet 5.5 cheaper than Sonnet 4.6 for the same work?

Usually, but by less than the list prices suggest. Sonnet 5.5 lists at $2 input versus $3, but it uses a tokenizer that produces about 30% more tokens for the same text, per Anthropic. Measure your own prompts with the token counting endpoint on both models to get the real ratio.

Does the Claude API charge more for long prompts?

Not on current models. Anthropic states that Claude 4.6 and later models bill the full 1M-token context window at standard per-token rates, so a 900,000-token request costs the same per token as a 9,000-token request. Long prompts still cost more in total because they contain more tokens.

Sources and further reading

  1. Anthropic pricing documentation (verified 2026-10-04)
  2. Anthropic prompt caching (verified 2026-10-03)

FrugalAI uses primary documentation and published research where possible. Product capabilities and prices can change; verify vendor details before procurement or production changes.