Trending: On-device modelsSearch
iHeartGeek
iTECH

Anthropic ships Claude Opus 5.5 at 40% less cost

Claude Opus 5.5 matches Claude Fable 5.1 on most work, undercuts Claude Opus 5 by 40% on typical workloads, and arrives with the tightened cyber and biology safeguards.

Anthropic’s Claude Opus 5.5 launch artwork: the model name in white serif type over a collage of a sunset horizon, a yellow notebook panel and a red rock face

Anthropic has released Claude Opus 5.5, the first model in a new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Claude Opus 5, and it is available from today in Claude’s own apps and through partners including GitHub Copilot.

Cheaper tokens, faster answers

The pricing is the headline. Input tokens are $4 per million and output tokens $20 per million, which is 20% below Opus 5, while cache reads - the part of the bill that dominates agentic and coding work - fall to $0.20 per million, 60% lower. Anthropic says Opus 5.5 also generates output more than 30% faster than its predecessor and needs less compute to serve, which is what makes the discount possible rather than promotional.

Subscription users get something less usual than a price cut: the five-hour usage limits rise on the Pro, Max, Team and seat-based Enterprise plans, and every subscription account receives a rate-limit reset it can bank and spend whenever it chooses. A stored reset is a small feature with an outsized effect on how people plan long jobs.

Safety testing and the cyber question

Anthropic tested the model before release with external evaluators, naming Frontier Design and METR, and reports that Opus 5.5 posts the best scores of any model it has tested on its automated behavioural audit - thousands of simulated scenarios that look for hard-to-reverse actions, work outside the boundaries it was given and susceptibility to prompt injection. The alignment suite has been widened to cover longer tasks, impossible tasks and scenarios modelled on real incidents, and the full evaluation sits in the model’s system card.

Because the model is comparable to Claude Mythos 5.1 in biology and cybersecurity, it ships with the safeguards built for Anthropic’s most capable models. Organisations can apply now to the Life Sciences Verification Program for biology work, and the Cyber Verification Program expands to verified security practitioners in the coming weeks. One detail is worth reading twice: where those safeguards intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier model development tasks by Claude Opus 5, so some published scores are a relay between generations.

How it scores

On Anthropic’s own tables Opus 5.5 leads Claude Fable 5.1 and Claude Opus 5 across agentic coding, computer use and knowledge work: 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0, 1,846 on GDPval-AA v2.1, 40.0% on AutomationBench, 67.7% on Humanity’s Last Exam with tools and 81.8% partial on OSWorld 2.0. The company adds its own caveat that at this level benchmark margins are a less reliable guide to real-world difference, and that the gap to Fable 5.1 is narrower in practice than the scores imply.

Where you can use it

Alongside Claude’s own products, GitHub shipped Opus 5.5 into Copilot on the same day for Pro+, Max, Business and Enterprise users, selectable in Visual Studio Code, Visual Studio, the Copilot CLI, the coding agent, the Copilot app, github.com and GitHub Mobile. GitHub notes that its outputs carry a watermark, that this changes neither quality nor cost, and that its own early testing resolved tasks comparably to Opus 5 using significantly fewer steps and tokens. Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks.

Our opinion

Anthropic spent early September asking the industry to pace the frontier, and has now shipped a model that matches its best frontier system on most work at a 40% discount. Those two positions are not contradictory, but they only make sense together if you accept that price is the strongest safety lever anyone actually holds: capability that gets cheaper gets used more widely, and a company that wants deployment decisions to be deliberate has to make the deliberate option affordable.

The disclosure that safeguarded cyber work was completed by Opus 4.8 and safeguarded biology work by Opus 5 is the most honest line in the announcement. It tells you that the headline numbers describe a system that hands the sharpest tasks to older models rather than a single agent that refuses them, and that the safeguard is a router as much as a refusal. Most vendors would have quietly published the merged score and let readers assume one model did all of it.