Home/Blog/Claude Opus 5
Model guide

Claude Opus 5,
and the dial that matters.

Opus 5 landed on 24 July 2026 at the same price as the model it replaces, which makes it an easy upgrade to wave through. Two of its changes will break working code, and one of them fails loudly while the other fails quietly.

Models · 5 min read

Opus 5 keeps Opus 4.8’s pricing at $5 per million input tokens and $25 per million output, with a one million token context window and up to 128,000 tokens of output. Anthropic describes it as a step change in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling, and it is now the default model on Max and the strongest option on Pro. The headline capability claim is that it converts extra effort into better results more reliably than any earlier Opus.

01

Thinking is on by default and effort is the control

On Opus 5 the model decides when and how much to think on each turn, and the effort parameter sets the depth. The available levels run low, medium, high, xhigh and max. Because the model converts effort into results more reliably than its predecessors, the level you choose now carries more weight than it used to, which cuts both ways: better answers when it matters, more tokens spent when it does not.

02

The breaking change that fails loudly

Disabling thinking is only accepted at effort high or below. Set thinking to disabled at xhigh or max and you get a 400 error. On Opus 4.8 those two settings were independent of each other, so any code carrying both a high effort level and a disabled-thinking flag stops working the moment you change the model ID. Better to find that in a test than in production at four on a Friday.

03

The trap that fails quietly

Your max_tokens value is a hard limit on total output, and total output now includes thinking as well as the response text. A workload that ran comfortably with thinking off on Opus 4.8 can start truncating on Opus 5 without any error, because the budget you set is being spent on reasoning before it reaches the answer. Anthropic flags this explicitly in its migration notes. Revisit the number.

04

Two smaller changes worth knowing

The prompt cache minimum drops from 1,024 tokens to 512, which brings caching into range for shorter system prompts that previously could not use it. There is also a fast mode in research preview on the API, running at higher speed for a higher price. Neither is headline material; both change what is worth optimising.

05

If you are on a subscription rather than the API

Opus 5 is the default on Max, so your model changed without you choosing it. Nothing is broken by that, but the thinking-on default means a given task can consume differently than it did a month ago. If your usage pattern shifted recently and you cannot explain why, this is the first place to look.

The comparison people will quote at you

Launch coverage put Opus 5 near Fable 5 performance at half the price, which is a fair summary and an incomplete one. It holds for general work and matters less for the narrow areas where the models are deliberately configured differently. As always, the answer that counts is what happens on your own tasks, which takes an afternoon to establish and saves months of arguing. That afternoon is more or less what a Claude Code coaching session is for.

See Claude Code coaching
Questions

Things people
usually ask.

Get started

The model got better.
Did your briefs?

A better model amplifies a good brief and a bad one equally. Bring your real work to a session with a specialist and find out which one you have been writing. We match you within 24 hours.

Book now →
Free · 5 minutes · no card · matched within 24 hours