Skip to content

[ SIGNAL: LIVE ]

Emergent Minds

>>_

Posts

View all →

A Dial Worth Turning: Claude Opus 5's Prose, and the Style Guide Anthropic Wrote Against Its Own Model

Opus 5 writes 510 words where Opus 4.5 wrote 158, with 2.3 times the em dashes and twice the 'load-bearing'. Arena measured it, Hacker News named it, Reddit downgraded over it. Anthropic's answer arrived in Fable 5.1's prompting docs: a paragraph defining 'mannered prose', with the model's own tics as the examples, for you to paste into your prompt. The vendor wrote the style guide against its own model, and shipped it as your job.

READ

The Thirty-Cent Judge: TypeSafe's Jev on a Real Product-Matching Queue

TypeSafe launched Jev on September 15 claiming 193x faster and 444x cheaper than frontier LLMs, zero hallucination, and calibrated probabilities. The launch evals measure agreement with GPT-6 and Fable 5.1, not correctness, and no calibration curve has been published. I had a better test in Pricogni: 9,081 low-confidence product matches a human review queue was never going to clear. One Noul, one Choice, 150 lines, 32 cents, 13 minutes. Half the queue was flankers, a fifth was publishable, and the one time it disagreed with a human reviewer the model was right.

READ

Careful, Not Thorough: Claude Opus 5.5 vs GPT-6 Sol on Real Code

Anthropic and OpenAI shipped Claude Opus 5.5 and GPT-6 Sol 101 minutes apart, and neither benchmarked the other. So I ran both on the same work: four self-contained tasks, then six real merged changes replayed on a large production monorepo, 20 runs each. Opus never broke a passing test. Sol did five times. Both covered the change equally well, and Sol cost 3.5x less.

READ