# Claude Sonnet 5.5's Max Effort Added Agents and Lowered Its Score

**Summary:** Claude Sonnet 5.5 promises lower task costs and 30% faster output, but Anthropic's own FrontierCode footnote shows Max effort adding scope and hurting the score. I would route effort by task instead of setting it globally.

- Canonical: https://markhuang.ai/news/claude-sonnet-5-5-max-effort-benchmark-worse
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-28
- Section: News
- Tags: Claude Sonnet 5.5, Claude Code, Coding Agents, Model Evaluation, AI Cost
- Source: [Anthropic](https://www.anthropic.com/claude-sonnet-5-5)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![A mechanical task line passes simple cards through a fast lane while a second lane loops through extra inspection arms](https://cdn.markhuang.ai/news/claude-sonnet-5-5-max-effort-benchmark-worse/hero.webp)

*More effort can buy more checking, but it can also add machinery that the job never needed.*

Anthropic [released Claude Sonnet 5.5](https://www.anthropic.com/claude-sonnet-5-5) on September 28, 2026, with an unusually practical pitch: output is more than 30% faster than Sonnet 5, while the cost per task can be up to 30% lower. The token rates have not moved. They remain $2 per million input tokens and $10 per million output tokens.

My read is that the model is less interesting as a blanket upgrade than as a routing choice. Anthropic says Sonnet 5.5 works best on well-scoped everyday jobs, while Opus 5.5 remains the choice for open-ended work that needs sustained judgment. The new Sonnet can narrow that gap when I turn effort up, but its own launch data shows why I should not leave the dial at maximum.

## The cheaper lane depends on the effort setting

Anthropic reports 70.6% for Sonnet 5.5 on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5. That is a vendor-reported result, not an independent verdict on every coding workflow. The cost curves are more useful to me. On several of Anthropic's tests, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for roughly one tenth of the cost per task. At higher effort, it can approach Opus 5.5 in both performance and cost.

This is the product decision hidden inside the benchmark page. Claude apps and Claude Code default to Medium effort, while the Claude Platform defaults to High. A team can therefore call the same model name and get a different balance of latency, token use, and checking depending on the surface and configuration.

That matters because the model Sonnet 5.5 replaces had a value problem. A July [independent analysis of Sonnet 5](https://fronset.ai/benchmark/journal/claude-sonnet-5-benchmark-analysis/) found it fast and capable, but recorded no cheapest-good-enough wins in a 52-task benchmark with partial model coverage. Anthropic is directly answering that complaint with fewer steps and fewer tokens. I still need to verify the answer on my own work.

## The best evidence is buried in a footnote

On FrontierCode, Sonnet 5.5 scored lower at Max effort than at Xhigh. Anthropic explains that Max effort made the model run Claude Code's code-review skill more often. That skill split the review among several subagents. In two cases examined by Cognition, the extra work caused either a timeout or edits beyond the task's scope, both of which lowered the score.

I find that detail more useful than the headline percentage. The model did not simply think longer. The higher setting changed the workflow, added workers, and created new ways to fail. A configuration that sounds safer can produce more review work or miss the finish line.

This is the same pattern I found when looking at [why stronger coding agents may need fewer tools](https://markhuang.ai/news/stronger-coding-agents-fewer-tools): the useful unit of evaluation is the model, harness, tools, and task together. The effort label is part of that system. It is not a quality slider I can push to the right without consequence.

> **Info:**
>
> I would start routine fixes at Medium, reserve High for tasks where another pass has clear value, and require evidence before enabling Xhigh or Max. The metric is cost per accepted result, including retries and human review.

## The API migration has its own bill

The [Sonnet 5.5 migration guide](https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide) makes clear that changing the model ID is only the first step. Adaptive thinking now runs when a request omits the `thinking` field. Thinking tokens count as output tokens, and code that assumes the first response block contains text can break because a thinking block may arrive first.

Developers who want no up-front thinking must use `between_tools` instead of the old disabled setting. That mode only works through High effort. Xhigh and Max require adaptive thinking, and incompatible combinations return a 400 error. Anthropic's own checklist tells teams to rerun their effort sweep and set a new cost baseline.

That is good migration advice because the advertised savings are averages from Anthropic's testing. A long tool loop, an oversized output budget, or a review skill that fans out across agents can spend the savings back. I would replay production-shaped tasks and record accepted outcomes, tokens, wall-clock time, timeouts, and scope corrections.

## Fallbacks make the model name less final

Sonnet 5.5 is also the first Sonnet launch with Anthropic's stronger cyber safeguards. Routine software work is meant to continue normally, but the company says higher-risk cybersecurity requests will visibly fall back to Sonnet 5. The [system card](https://www.anthropic.com/claude-sonnet-5-5-system-card) provides the safety evaluation behind the release, while Anthropic says its automated behavioral audit covered roughly 1,850 scenarios.

A fallback may be the right safety choice, but it changes what a production evaluation measures. If a request begins on Sonnet 5.5 and completes on Sonnet 5, the log needs to preserve that fact. Otherwise a team may attribute latency, refusal behavior, or task quality to the wrong model.

The first [public launch thread](https://www.reddit.com/r/ClaudeAI/comments/1wslxzs/introducing_claude_sonnet_55_the_second_model_in/) was mostly too early and too anecdotal to judge quality. It did surface immediate concern about cyber fallbacks, alongside excitement about the benchmark gap. I would treat both reactions as test ideas, not evidence that the model is either crippled or transformative.

## I would buy the efficiency, then cap it

Sonnet 5.5 appears designed to solve a real problem: Sonnet 5 could be fast without being the cheapest model that cleared a task's quality bar. Anthropic's new release claims better results with fewer tool calls, fewer tokens, and shorter waits, all at the same token price.

I would pilot that claim in the lane Anthropic describes: bounded coding tasks, document work, and repeatable jobs where speed matters. I would keep Opus for work where the hard part is deciding what the task should become, and I would not use Max as a global default. One of Anthropic's own benchmarks already shows the failure mode. The model worked harder, recruited more help, and produced a worse scored result.
