Every frontier model now has an effort setting, from low to max. Claude Code runs both Claude 5.5 models at medium by default. The Claude API runs Sonnet 5.5 at high. Launch tables quote scores at xhigh or max, and say more effort buys a better score.
This week I ran Opus 5.5 against GPT-6 Sol, then Sonnet 5.5, on real changes at medium. Then I ran all three at every other level. On real work, the best Opus setting was the default, and every level above it broke more.
Real Work, Not a Benchmark
No puzzles and no synthetic tasks: six changes that were actually merged into a large production TypeScript monorepo. Each run started from the commit before the change, with a ticket-style prompt and no git history. The tests the original fix shipped with did the grading. 20 runs per model per level, 300 in all, at API list prices. No run timed out, fell back to another model, or hit a rate limit.
The Opus and Sol medium runs first ran on September 23. I reran them alongside the other levels. Opus reproduced its result exactly, change by change.
The Results
A clean run passes every test, old and new.
| 20 runs each | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| Opus 5.5 clean runs | 11 | 13 | 12 | 12 | 11 |
| Broke a passing test | 0 | 0 | 1 | 1 | 2 |
| Cost | $16.97 | $44.10 | $65.16 | $139.79 | $307.16 |
| Sonnet 5.5 clean runs | 10 | 11 | 11 | 10 | 13 |
| Broke a passing test | 0 | 0 | 0 | 0 | 0 |
| Cost | $8.38 | $11.76 | $26.60 | $96.04 | $415.18 |
| GPT-6 Sol clean runs | 9 | 9 | 7 | 6 | 7 |
| Broke a passing test | 5 | 5 | 7 | 6 | 4 |
| Cost | $9.22 | $13.71 | $17.49 | $22.37 | $30.09 |
- Opus peaks at medium. Below it, Opus missed more: at low it solved an edge-case-heavy change once in three runs, against three at medium. Above it, Opus broke more: one run at high and xhigh, two at max.
- Effort is not care. Sol broke tests at every level. Sonnet broke nothing at any level. Opus was the only model that moved, and above medium it moved the wrong way.
- One setting beat its own medium. Sonnet at max got 13 clean runs against 11, all from one edge-case change it had missed at medium. It cost 35x as much, and more than Opus at max: $415 against $307.
Every Opus regression was the same one. It rewrote an existing user-facing message on Change A and dropped the instruction the message carried. Sol had the same habit, and effort made it worse for both:
| Runs that dropped the instruction (of 4) | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| Sonnet 5.5 | 0 | 0 | 0 | 0 | 0 |
| Opus 5.5 | 0 | 0 | 1 | 1 | 2 |
| GPT-6 Sol | 1 | 2 | 4 | 4 | 4 |
Real Work vs Benchmarks
I don’t think the launch tables are wrong. They measure something else. This is my reading of the data, not a proven cause, and it fits what I argued when Berkeley gamed the leaderboards.
- Benchmarks grade whether the task got done. On my bench, that barely moved: every model fixed 83% to 94% of the new tests at every level. The differences were at the edges: breaking something that already worked, or missing a case the ticket didn’t spell out.
- Real changes land in code that already works. Many benchmark tasks are self-contained, with little existing work to break. My changes landed in about 3,700 files. More effort meant bigger diffs: Opus’s Change A diff grew from about 480 added lines at medium to about 830 at max. More rewriting meant more chances to drop something.
- Effort helps only when the gap is reasoning depth. Sonnet’s one gain at max was a set of edge cases, the kind of problem benchmarks are built from.
- Effort doesn’t fix a reading of the ticket. On the widest change, 26 files, no model produced a clean run at any level. At medium and high, all three skipped the same part, one the ticket named.
- The benchmark closest to real work agrees. FrontierCode penalizes out-of-scope changes, and it is the one Anthropic benchmark where Sonnet scored lower at max than at xhigh. Anthropic’s footnote blames out-of-scope edits. My bench didn’t reproduce that for Sonnet, but it saw the same failure in Opus.
Effort changed how much each model wrote and how far it searched. It never made a model more careful. On a subscription that meters in API dollars, it also sets how fast the plan runs out.
That last point is no longer hypothetical. This week Thibault Sottiaux of OpenAI’s Codex team announced that the reopened $200 Pro plan will “net out at half the dollar in API spend compared to the old Pro $200 plan”. Opus at max used 7x the list-price value of Opus at medium, for two fewer clean runs.
— Simon Willison, on Opus 5.5 at max efforteffectively useless
Max was not useless on my bench. It was the best Sonnet setting, the worst Opus setting for regressions, and no help to Sol.
Low Is Underrated
- Sonnet at low is the cheapest careful setting. 10 clean runs, zero regressions, $8.38.
- Opus at low got 11 clean runs and zero regressions for 38% of its medium bill.
- Nothing skipped its checks. Anthropic’s guide warns that Sonnet 5.5 at low “sometimes skips a check that exercises the change”. All 60 low runs ran tests, a typecheck or a build before finishing.
What This Doesn’t Show
- Three or four runs per change. Most gaps in the table are one or two runs. Sonnet’s gain at max is three runs on one change.
- Claude effort is inferred. The CLI accepted each level and thinking tokens moved with it, but the requests don’t record the setting. The Codex logs confirm Sol’s.
- One codebase, my prompts. A different codebase or a vaguer ticket could move every number.
- List prices, not my bill. The 300 runs came to $1,224 at list price, on subscriptions.
What I’m Doing With It
- Leave Opus 5.5 on medium. Best Opus setting on both days, with no regressions.
- Sonnet 5.5 at low or medium for cheap work with good tests. Zero regressions at both.
- Max is a retry, not a setting. When a medium run fails the tests, a Sonnet max retry is a reasonable second attempt.
- Don’t raise effort to fix carelessness. A test or a review step catches a dropped instruction. More thinking made it more likely.
The effort setting looks like a quality dial. On real work it behaved like a spend dial. The best setting for Opus was the one it ships with.



