Opus 5 writes 510 words where Opus 4.5 wrote 158, with 2.3 times the em dashes and twice the 'load-bearing'. Arena measured it, Hacker News named it, Reddit downgraded over it. Anthropic's answer arrived in Fable 5.1's prompting docs: a paragraph defining 'mannered prose', with the model's own tics as the examples, for you to paste into your prompt. The vendor wrote the style guide against its own model, and shipped it as your job.
Read more →Anthropic and OpenAI shipped Claude Opus 5.5 and GPT-6 Sol 101 minutes apart, and neither benchmarked the other. So I ran both on the same work: four self-contained tasks, then six real merged changes replayed on a large production monorepo, 20 runs each. Opus never broke a passing test. Sol did five times. Both covered the change equally well, and Sol cost 3.5x less.
Read more →Boris Cherny, who built Claude Code, says engineering, product, design and data science are melting into one role, and what's left is five archetypes: Prototyper, Builder, Sweeper, Grower, Maintainer. I read the list and realised I'm all five, because building solo with agents leaves no one to hand a phase to. The framework is thirty years old. What's new is that it just became the primary axis instead of the secondary one.
Read more →TypeSafe's Jev is an AI model that doesn't write. It answers questions with a yes/no probability, a pick or a score. I gave it 9,081 product matches a human review queue was never going to clear. It judged them all for 32 cents in 13 minutes. Eight days later, checked against barcodes, it got the job.
Read more →Benchmarks say turn effort up. On real work, it went the other way. I ran Claude Opus 5.5, Claude Sonnet 5.5 and GPT-6 Sol at low, medium, high, xhigh and max effort on six real changes from a production monorepo: 300 runs. Opus peaked at medium, the Claude Code default, and reproduced it exactly on a second day. Above medium it started breaking tests. More effort never made any model more careful.
Read more →From 20 lines of shell to production apps. Anthropic renamed Claude Code SDK to Agent SDK because deep research is now a first-class use case.
Read more →OpenAI's GPT-6 Astra reached every paid ChatGPT plan, the API, Copilot and OpenRouter 27 hours after launch, at Fable 5.1's exact price. It sits two points behind Fable on the independent index at 42% of the cost per task, refuses exploit-writing by default, and hides its reasoning. I ran the same code review at low, high and max effort. Low found five real bugs in 60 seconds. Max found seven in seven minutes and got the ranking right.
Read more →Boris Cherny shared his workflow for the tool he built. The setup is surprisingly vanilla. The philosophy is worth studying.
Read more →The official Claude Code plugin that lets agents work autonomously for hours. When to use it, when not to, and the philosophy behind letting AI fail repeatedly until it succeeds.
Read more →Claude Code has no profile switcher. An account is one Keychain entry keyed by a directory hash, so three accounts can share one set of plugins, skills and transcripts by symlink. Only the prompt history refuses. The usage endpoint has a model-scoped weekly bucket and rate-limits polling. I packaged it as ccseats.
Read more →