
Anthropic launched Claude Opus 5.5 on September 22, 2026, the first model in its Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. Anthropic also says the model was tested before release by external evaluators including Frontier Design and METR.
The release focuses on long-running, cross-file work that requires multiple tool steps rather than only single-turn answers. Anthropic’s internal and early-access testing covers agentic coding, computer use, knowledge work, and business workflows. Most of the results on the launch page are still vendor, partner, or early-tester evidence, so they should not be treated as independent estimates of performance in every production environment.
For coding, Anthropic positions Opus 5.5 as particularly capable on large codebase migrations and audits. One early tester reported completing a 680,000-line migration in less than a day; another test had the model improve page-load times across a web application and succeed 39 times out of 40. These examples suggest stronger long-horizon task execution, but engineering teams still need to review changes, run tests, and maintain a rollback path.
The pricing details are relevant to agent workloads: $4 per million input tokens, $20 per million output tokens, and $0.20 per million cache-read tokens. Anthropic says a typical workload costs 40% less than Opus 5 at default settings and that output is more than 30% faster. For an agent that runs for hours, the number of steps and the effectiveness of context caching can matter more than the headline token price.
On safety, Anthropic says Opus 5.5 achieved its best result to date on the company’s automated behavioral audit, with expanded testing for longer tasks, impossible tasks, and scenarios modeled on real incidents. It also reports stronger resistance to prompt injection than Opus 5 and describes a classifier before each action, an auditable sandbox, and pre-merge code review as part of the deployment approach. These remain product claims; a full assessment requires the system card, test conditions, and external replication.
The benchmark table includes Terminal-Bench, FrontierCode, CursorBench, GDPval, and OSWorld. Anthropic cautions that benchmark margins are becoming a less reliable guide to real-world differences at this capability level. Effort settings, tools, safety interventions, and harness details all affect results, so a serious evaluation should track task completion, tool-call count, human rework, and the cost of mistakes together.
Opus 5.5 is available across several plans, with Anthropic also raising five-hour usage limits for some subscriptions and promising Sonnet 5.5 and Haiku 5.5 in the coming weeks. Access for life-science and cybersecurity work is being handled through verification programs. As capability rises, identity, permissions, and high-risk-domain review become part of the product design rather than an afterthought.
The broader signal is that coding agents are moving from short-task assistants toward work units that can run for extended periods. Once a model can retain context, edit many files, and verify its own work, evaluation has to include permission boundaries, action traceability, approval points, and safe recovery. Opus 5.5’s figures are worth testing, but controlled pilots and measurable engineering outcomes should determine adoption.



