Anthropic released Claude Sonnet 5.5 at 11:03 AM PT on September 28. Six days earlier I ran Opus 5.5 against GPT-6 Sol on six real changes from a production monorepo. Opus was the careful one: zero regressions in 20 runs. Sol cost 3.5x less and broke a passing test five times.

Sonnet 5.5 has exactly Sol’s per-token price. So the question is simple: can you get Opus’s care at Sol’s price?

The Sticker

Per million tokensInputCached inputOutput
Claude Sonnet 5.5$2$0.20$10
GPT-6 Sol$2$0.20$10
Claude Opus 5.5$4$0.20$20

Anthropic’s developer guide splits the work cleanly. Sonnet 5.5 is for “well-scoped everyday tasks”, Opus 5.5 for “complex work requiring careful judgment”.

For the hardest long-horizon work, an Opus model is the better choice.

— Anthropic, Sonnet 5.5 prompting guide

Anthropic’s own table is less clear about it. Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0 (70.6% against 66.4%) and trails it on every other row, by up to 8 points. The system card explains part of that win. Safeguards sent 1.2% of Sonnet’s Terminal-Bench requests to a fallback model, affecting 1.5% of trials. For Opus the figure was 10% of trials. The stated standard error is 2.5 points.

The Same Six Changes

I used the same bench as last week: six merged changes, each replayed from the commit before it on a clean snapshot with no git history, and graded by the tests the original fix shipped with. Same ticket-style prompts, same global rules file, medium effort, 20 runs. Every Sonnet run was served by Sonnet 5.5. None fell back.

20 runs eachSonnet 5.5Opus 5.5GPT-6 Sol
Runs that passed every test11138
Runs that broke a test that passed before005
New tests fixed88%93%90%
Original-fix files also changed82%85%84%
Output tokens247k495k192k
Total cost at list price$11.76$45.81$13.16

Per change, clean runs and mean cost per run:

ChangeSonnet 5.5Opus 5.5GPT-6 Sol
A: 15 files, 1 workspace4/4, $0.334/4, $1.120/4, $0.44
B: 17 files, 3 workspaces1/4, $0.890/4, $3.361/4, $0.79
C: 26 files, 3 workspaces0/3, $1.010/3, $4.760/3, $0.78
D: 15 files, 1 workspace0/3, $0.353/3, $0.981/3, $0.63
E: 11 files, 3 workspaces3/3, $0.403/3, $1.703/3, $0.76
F: 13 files, 2 workspaces3/3, $0.533/3, $1.853/3, $0.58
  • It keeps what works. On Change A, Sol rewrote an existing user-facing message in all four runs and dropped the instruction it carried. Sonnet kept it in all four, like Opus. It also wrote smaller diffs than either: about 184 added lines on A, against about 240 for Sol and 460 to 495 for Opus.
  • One change is almost the whole gap. On Change D, all three Sonnet runs failed the same six tests: the edge cases where the input should be ignored or left alone. Sonnet did add the safety check. It handled the edges differently from the original fix, and it did so the same way every time. On a different codebase, a different change would sit in that slot.
  • Coverage is slightly behind. Sonnet changed 82% of the files the original fix changed. On Change B, one screen that no test covers needed an update: Opus fixed it in three of four runs, Sonnet in one, Sol in none. On Change C, all three models skipped the same part of the codebase.
  • The price gap is real. Opus cost 2.8x to 4.7x as much as Sonnet per change, 3.9x overall. Sonnet came in under Sol overall, although Sol’s figure has no cache-write charge and Sonnet’s does.
Care came down a tier

Last week’s result was that Opus buys fewer regressions, not more coverage. Sonnet 5.5 keeps the fewer regressions at Sol’s price. What it gives up against Opus is a few percent of coverage and one consistent blind spot.

Only Cheap at Medium

The cheap-model story depends on the effort setting. Artificial Analysis runs both models at max effort. Sonnet 5.5 writes 410M output tokens across its index, against 260M for Opus 5.5, and costs $7.60 per task, against $5.98. At max, the cheaper model is the more expensive one.

More thinking does not buy more score either. On Anthropic’s own FrontierCode table, Sonnet scores 52.1% at xhigh and 46.2% at max: at max it more often ran a code-review skill and timed out or made out-of-scope edits. Simon Willison’s max-effort pelican “thought for 128,000 tokens” and ran out before it produced an SVG. At xhigh it took 41 seconds and about 6 cents.

Claude Code runs Sonnet 5.5 at medium by default. The API defaults to high. Check which one you are paying for.

Update: High Effort

After publishing, I reran all three models at high effort on the same six changes, 20 runs each. The Codex logs confirm Sol ran at high. For the Claude runs, the CLI accepted the flag and output tokens doubled, but the requests don’t record the setting.

20 runs each, medium / highSonnet 5.5Opus 5.5GPT-6 Sol
Runs that passed every test11 / 1113 / 128 / 7
Runs that broke a test that passed before0 / 00 / 15 / 7
New tests fixed88% / 90%93% / 94%90% / 89%
Original-fix files also changed82% / 85%85% / 88%84% / 85%
Total cost at list price$11.76 / $26.60$45.81 / $65.16$13.16 / $17.49
  • High buys coverage, not care. No model gained a clean run. Coverage rose by one to three points. On the untested Change B screen, Sonnet went from one run in four to four in four. Opus fixed it in four, Sol in none at either level.
  • Sonnet still broke nothing. That makes zero regressions in 40 runs. One Opus run at high dropped the same Change A instruction that Sol drops. Sol dropped it in all four runs again.
  • Change D closed a little. One of three Sonnet runs at high passed every test. The other two failed the same six edge cases.
  • Sonnet pays most for it. High cost Sonnet 2.3x its medium bill, against 1.4x for Opus and 1.3x for Sol. At high, Sonnet costs more than Sol. It matches Opus at medium on coverage, at 58% of the price.

Was Opus Nerfed?

On launch day a friend said Opus 5.5 “seems token hungry today”. A few Reddit posts say the same about Max plan limits. About as many say usage is fine, and I found nothing from Anthropic about a change.

My own Claude Code transcripts don’t show it. Across 888 Opus 5.5 turns after the launch, the median output per turn was 368 tokens, against 355 earlier that day. Context per turn was flat at about 275k tokens. That measures the model, not the meter. One Reddit user thinks cache reads now count against the Max quota. If that is true, tokens would stay flat while limits drain faster. I can’t confirm it from my data, and neither can anyone else yet.

What This Doesn’t Show

  • Three or four runs per change. Enough to see a habit repeat, not tight rates. Sonnet’s gap to Opus is one change.
  • Medium and high only. I didn’t test xhigh or max on my own changes.
  • One codebase, my prompts, a newer CLI. The Opus and Sol runs used an earlier Claude Code version than the Sonnet runs.
  • List prices, not my bill. All of it ran on subscriptions.

What I’m Changing

  • Sonnet 5.5 at medium replaces Sol where tests and review are strong. Same price, and no regressions in 40 runs across two effort levels.
  • Medium, not high, for Sonnet. High more than doubled the bill and added no clean runs.
  • Opus 5.5 stays for thin tests and unattended runs. Its edge over Sonnet is small, and it is the edge-case handling Sonnet missed on Change D.
  • Never max on Sonnet. At max it costs more than Opus and scores lower than at xhigh.

Last week the choice was care or price. This week the care costs the same as the cheap option, as long as you leave the effort on medium.