Trending Posts

Trending

Most popular posts from the last 7 days

10 trending posts

A Dial Worth Turning: Claude Opus 5's Prose, and the Style Guide Anthropic Wrote Against Its Own Model

Opus 5 writes 510 words where Opus 4.5 wrote 158, with 2.3 times the em dashes and twice the 'load-bearing'. Arena measured it, Hacker News named it, Reddit downgraded over it. Anthropic's answer arrived in Fable 5.1's prompting docs: a paragraph defining 'mannered prose', with the model's own tics as the examples, for you to paste into your prompt. The vendor wrote the style guide against its own model, and shipped it as your job.

Read more →

Careful, Not Thorough: Claude Opus 5.5 vs GPT-6 Sol on Real Code

Anthropic and OpenAI shipped Claude Opus 5.5 and GPT-6 Sol 101 minutes apart, and neither benchmarked the other. So I ran both on the same work: four self-contained tasks, then six real merged changes replayed on a large production monorepo, 20 runs each. Opus never broke a passing test. Sol did five times. Both covered the change equally well, and Sol cost 3.5x less.

Read more →

The Archetype Under the Title

Boris Cherny, who built Claude Code, says engineering, product, design and data science are melting into one role, and what's left is five archetypes: Prototyper, Builder, Sweeper, Grower, Maintainer. I read the list and realised I'm all five, because building solo with agents leaves no one to hand a phase to. The framework is thirty years old. What's new is that it just became the primary axis instead of the secondary one.

Read more →

The Default Was Right: Opus 5.5, Sonnet 5.5 and GPT-6 Sol Effort Levels on Real Work

Benchmarks say turn effort up. On real work, it went the other way. I ran Claude Opus 5.5, Claude Sonnet 5.5 and GPT-6 Sol at low, medium, high, xhigh and max effort on six real changes from a production monorepo: 300 runs. Opus peaked at medium, the Claude Code default, and reproduced it exactly on a second day. Above medium it started breaking tests. More effort never made any model more careful.

Read more →

GPT-6 Astra Is On Every Plan: What It Costs, What It's Good At, and Which Effort Level to Use

OpenAI's GPT-6 Astra reached every paid ChatGPT plan, the API, Copilot and OpenRouter 27 hours after launch, at Fable 5.1's exact price. It sits two points behind Fable on the independent index at 42% of the cost per task, refuses exploit-writing by default, and hides its reasoning. I ran the same code review at low, high and max effort. Low found five real bugs in 60 seconds. Max found seven in seven minutes and got the ranking right.

Read more →