The Default Was Right: Opus 5.5, Sonnet 5.5 and GPT-6 Sol Effort Levels on Real Work
Benchmarks say turn effort up. On real work, it went the other way. I ran Claude Opus 5.5, Claude Sonnet 5.5 and GPT-6 Sol at low, medium, high, xhigh and max effort on six real changes from a production monorepo: 300 runs. Opus peaked at medium, the Claude Code default, and reproduced it exactly on a second day. Above medium it started breaking tests. More effort never made any model more careful.
READ



