Guides 9 min read

Claude Opus 5.5 for Agents: Real Cost and the Rerouting Problem

Opus 5.5 is 40% cheaper than Opus 5. It also silently hands some agent calls to Opus 4.8. Real cost per task, and how to detect the swap.

Shabnam Katoch

Shabnam Katoch

Growth Head

Model economics: Claude Opus 5.5 for Agents, real cost and the rerouting problem
Free forever

Your agent. Working. Not broken.

One AI agent that just works.

No silent failures. Free forever, not a trial.

Start free

No credit card · No Docker · No config files

Anthropic cut Opus prices by 20% and cache reads by 60% on 22 September. It also shipped a safety layer that can answer your agent's request with a different, older model and not tell you in the response body. Here's what Opus 5.5 actually costs per agent task, and how to find out which model really answered.

You set your agent to claude-opus-5-5. It ran overnight. In the morning the work looks fine, the bill looks lower than last month, and you move on.

Then you open the raw responses and find claude-opus-4-8 sitting in the model field on a handful of calls. Nothing errored. Nothing warned you. A different model just did some of the work.

That isn't a bug, and it isn't a rumour. Anthropic published it alongside the model.

What Opus 5.5 actually costs

Anthropic released Claude Opus 5.5 on 22 September 2026, the first model in the Claude 5.5 family. The price cut is real and it's bigger than the headline suggests for anyone running agents.

Per million tokensOpus 5.5Opus 5Sonnet 5
Input$4$5$2
Output$20$25$10
Cache reads$0.20$0.50$0.20
Cache writes$5$6.25–
Fast mode$8 in / $40 out––

Prices from Anthropic's launch, checked 25 September 2026.

The line that matters for agents is cache reads. They fell 60%, from $0.50 to $0.20, and Anthropic says cache reads make up the majority of cost in agentic and coding work. That's why the company puts the overall saving at about 40% on typical workloads while the sticker price only dropped 20%.

Two more things changed. Output generates more than 30% faster, and Anthropic says the model uses fewer tokens per task at default settings, so the same job bills less twice over. There's also a fast mode in Claude Code and the Claude Platform at double the token rates, up to 2.5 times quicker.

The headline is not the saving: agent calls are mostly cache reads (instructions, tool schemas, history) with a sliver of fresh input. The sticker price fell 20% from $5 to $4, while cache reads fell 60% from $0.50 to $0.20, about 40% off a typical agent workload

The headline is a 20% price cut. The number that actually moves an agent's bill is the 60% cache-read cut, because your agent re-sends the same context on every single call.

Cost per agent task, worked out

Take a standard agent run: roughly 40,000 tokens of cached context (your instructions, tool schemas and history), about 3,000 tokens of fresh input, and 2,000 tokens of output.

Opus 5.5Opus 5Sonnet 5
Cache read (40K)$0.008$0.020$0.008
Fresh input (3K)$0.012$0.015$0.006
Output (2K)$0.040$0.050$0.020
Per task$0.060$0.085$0.034
At 200 runs a day$360/mo$510/mo$204/mo

So Opus 5.5 is about 29% cheaper than Opus 5 on this shape of work, and about 1.8 times the price of Sonnet 5. If Anthropic's fewer-tokens-per-task claim holds on your workload, the real gap against Opus 5 is wider than the table shows.

For comparison on the same day, OpenAI cut GPT-6 Sol to $2 and $10 per million, matching Sonnet 5's sticker (Opus 5.5 vs GPT-6 Sol vs Grok 4.7 runs the same task maths across all three). The frontier tier is getting cheaper fast, which is exactly why per-task maths beats per-token maths. Our LLM pricing guide keeps the full grid current.

One agent run on three models, with 40K cached, 3K input and 2K output tokens: Opus 5.5 $0.060, Opus 5 $0.085, Sonnet 5 $0.034. Cache reads move the bill, not the sticker

The rerouting problem: your agent might not be running the model you picked

Here's the part that deserves an architecture review rather than a shrug.

Opus 5.5 is the first Opus model to ship with the same class of safety classifiers Anthropic already runs on Fable 5.1, covering three areas: cybersecurity, biology and life sciences, and distillation (attempts to extract the model's behaviour to train another model).

When one of those classifiers fires, your request is not refused. Anthropic's own wording is that the safeguards "fall back to another model transparently":

  • Most flagged cybersecurity requests are handled by Opus 4.8
  • Flagged biology and frontier LLM development requests are handled by Opus 5

Your API call still succeeds. You get a normal response. A different model wrote it.

Anthropic is unusually candid about why: it says Opus 5.5 has "extremely strong cyber capabilities", and it benchmarked the model with production safeguards switched on, noting the fallbacks "likely reduce Claude Opus 5.5's performance on these benchmarks". The company is telling you its own scores understate the model because some of the test work was done by an older one.

For a chatbot, this is a footnote. For an agent, it isn't. A six-step chain where step four trips the classifier is a chain where one step ran on a model with different capabilities, different tool-calling behaviour and a different training cutoff. Your evals won't catch it either, because average behaviour stays fine while individual steps get handled elsewhere.

We saw how loud this can get with Fable 5. One developer running routine defensive threat-intelligence work in Claude Code logged 18 fallback events in three days, and on a single day 2,746 of 3,427 assistant messages in the main session came from Opus 4.8 rather than the model they'd configured. The flagged work was monitoring published CVEs and summarising public threat reports. No exploit development anywhere in it.

That's the shape of the risk: not malicious prompts getting blocked, but ordinary security-adjacent work quietly getting a different model. If your agent reviews dependencies, triages vulnerability reports, or reads security advisories, you are in the blast radius, and it belongs in your agent security review.

The safeguard isn't "your request was refused". It's "your request was answered by something else". Those need very different handling in code.

How to detect it

One field. Every response carries the model that actually generated it, and on a rerouted call it won't match what you asked for:

resp = client.messages.create(model="claude-opus-5-5", ...)

if resp.model != "claude-opus-5-5":
    log.warning("rerouted to %s", resp.model)   # e.g. claude-opus-4-8

Log that field on every call, today, before you do anything else. It costs nothing and it turns an invisible problem into a number you can look at. Also keep handling stop_reason: "refusal", because the fallback model can refuse too, and then you get a refusal after the swap.

One step, a different model: in a six-step agent chain on opus-5-5, a classifier fires at step 4 and opus-4-8 answers instead, while the response still returns 200 with response.model set to opus-4-8. Log the model field on every call

How to handle it

Three moves, in order of how much they help.

Split the agent. If one agent does dependency review and another drafts customer emails, don't run both on Opus 5.5 and hope. Put security-adjacent work on a model that won't reroute, and keep Opus 5.5 for the work where it's strongest.

Pin behaviour per step, not per agent. The chain is what breaks, so the fix belongs at the step level. Any step whose output feeds a later step should either be verified as non-flagging or run somewhere predictable.

Alert on the rate, not the event. One fallback a week is noise. Eighteen in three days is a broken workload, and the Fable 5 case shows how quickly it escalates from one to most of your traffic.

On BetterClaw, model choice is per agent rather than per account, so the security agent and the email agent can sit on different models without running two platforms. Bring your own Anthropic key and you see the real provider bill, including the calls that came back from a different model. Free plan, no card.

Opus 5.5 vs Sonnet 5: when is 1.8x worth it?

The cheaper Opus makes this a closer call than it used to be, but the answer hasn't flipped.

Use Sonnet 5 for the work an agent does fifty times a day: classification, triage, drafting, summarising, routine tool calls. At $0.034 a task against $0.060, you're paying 76% more for output most people can't distinguish in a blind read.

Use Opus 5.5 where the task is long, multi-step and expensive to get wrong. Anthropic reports Opus 5.5 at default effort beating Opus 5 at max effort on Terminal-Bench 4.0 for about a fifth of the cost, and matching GPT-6 Astra at roughly 40% of the price. One customer's C-to-Rust port of HAProxy finished in 9.5 hours on Opus 5.5 versus 12 on Fable 5.1, at 51% lower cost. That's the profile: big refactors, long agentic coding runs, work where a wrong answer costs an afternoon.

Match the task to the model: triage and drafting on Sonnet 5 at $0.034, long coding runs on Opus 5.5 at $0.060, and security research, which may reroute, on a model with no classifier layer. Route by task, not by favourite model

Use neither if the job is security research. That's the one place where the model you chose may not be the model you get.

The routing rule we'd actually write: Sonnet 5 as the default, Opus 5.5 for long multi-step coding, and anything security-adjacent pinned to a model without the classifier layer. Our model routing setup guide covers the escalation logic, and if cache reads now dominate your bill, prompt caching for agents is where the remaining money is.

Two smaller things worth knowing before you migrate. Thinking cannot be switched off on Opus 5.5, so any code that sets thinking: disabled needs revisiting. And Anthropic removed the five-hour usage caps for Pro, Max, Team and seat-based Enterprise subscribers in the same announcement, which partly walks back the limits story from earlier this month, covered in our Claude Code rate limit guide.

What this release actually tells you

Opus 5.5 is a genuinely good deal. Fable-class quality on most work, 40% off typical agent workloads, faster output. If you're on Opus 5, migrating is close to a free upgrade.

But the interesting thing isn't the price. It's that "which model am I running" has stopped being a question you answer once in a config file. Safety layers now route requests at runtime, and the answer can change per call without changing your code. That's a new category of thing to monitor, alongside latency and spend.

The teams who handle it well won't be the ones who picked the best model. They'll be the ones who logged the model field.

If you'd rather set models per agent than per account, and see exactly what each one costs, that's what we built. Free plan with 1 agent and 100 credits a month, your own API keys with no markup, then Basic at $19, Pro at $49 and Business at $149 a month. Start free or see the full pricing.

Frequently Asked Questions

What is Claude Opus 5.5 and how much does it cost?

Claude Opus 5.5 is Anthropic's model released on 22 September 2026, the first in the Claude 5.5 family. It costs $4 per million input tokens and $20 per million output, 20% below Opus 5, with cache reads down 60% to $0.20 and cache writes at $5. A fast mode in Claude Code and the Claude Platform runs at $8 and $40. Anthropic puts the total saving at about 40% on typical workloads.

How does Opus 5.5 compare to Opus 5 for agent work?

It's cheaper on every line and faster. On a typical agent task with 40,000 cached tokens, 3,000 fresh input and 2,000 output, Opus 5.5 works out around $0.060 against $0.085 for Opus 5, roughly 29% less. Anthropic also says Opus 5.5 uses fewer tokens per task at default settings and reports it beating Opus 5 at max effort on Terminal-Bench 4.0 for about a fifth of the cost.

How do I tell if my request was rerouted to Opus 4.8?

Check the model field on the response. If you requested claude-opus-5-5 and the response says claude-opus-4-8, a classifier fired and an older model answered. Log that field on every call and alert on the rate rather than individual events. Keep handling stop_reason: "refusal" too, because the fallback model can still refuse.

Is Opus 5.5 worth 1.8x the price of Sonnet 5 for agents?

For long, multi-step coding work, usually yes, because finishing a big refactor correctly is worth more than the token difference. For the jobs an agent does fifty times a day, such as triage, classification and drafting, usually no: Sonnet 5 runs about $0.034 a task against $0.060 and produces output most people can't tell apart. Route by task rather than picking one model for everything.

Is the rerouting a safety problem or a reliability problem?

Both, but the reliability side is what will bite you first. Anthropic designed the fallback so requests get answered instead of refused, which is friendlier behaviour, and it says the safeguards even cost it benchmark points. The risk for agents is that one step in a chain can run on a model with different capabilities and nothing in the response announces it, so multi-step workflows can degrade in ways your evals average away.

Want to skip the setup?

BetterClaw does this in 60 seconds. No Docker, no config files.

Start free
Tags:claude opus 5.5opus 5.5 pricingopus 5.5 vs opus 5opus 5.5 agentclaude opus 5.5 cost per taskopus 5.5 reroutingopus 5.5 vs sonnet 5opus 5.5 cache pricing
Share this article
Was this helpful?