In July I wrote that Grok 4.5 trained on the answer key: an earlier snapshot of the Cursor codebase, the thing CursorBench grades against, turned up in the training data, and Cursor pulled the scores. Five weeks later, Grok 4.6’s launch page does something I didn’t expect. It publishes a ten-row benchmark table that its rival wins.
Count them. The unbranded “Fable 5 Max” column takes six of ten rows, including the Artificial Analysis Intelligence Index (62 to Grok’s 61, with Opus 5 at 63 above both). DeepSWE, which Grok loses to GPT-5.6 Sol by 7.1 points, is on the hero chart. And CursorBench is back: reinstated after the contamination fix, scored clean at 69.9%, and losing to Fable’s 70.5%. The benchmark that got them caught, retaken honestly, and lost.
The wins are real too: the GDPVal-AA knowledge-work eval, AA-Briefcase, and a genuinely strange 15.8% on the Harvey legal benchmark against Sol’s 2.5%. At $2 in and $6 out per million tokens, the price argument from July still stands.
The Two Tells
Forced honesty has a residue, and you can see it in what the page still won’t do.
- The winner has no name. The post compares against “GPT-5.6 Sol” in prose, by name. The column that beats it six times is labelled “Fable 5 Max”, and the words Claude, Opus, and Anthropic appear nowhere on the page. Printing the loss is apparently easier than printing who inflicted it.
- One benchmark vanished. SWE-bench Pro, the stage for July’s token-efficiency misdirection, is absent from the page entirely. Every other benchmark from the 4.5 story came back. The one that anchored the misleading hero chart did not.
xAI reports Terminal-Bench v3.0 at 26%. Artificial Analysis reports Terminal-Bench v2.1 at 88.4%. Same benchmark family, different versions, wildly different scales. Any comparison that drops the version number is accidentally lying.
The Street Price
The table says tied with Sol. The first-hand reports on Hacker News, a day or two into real use, say something else.
— bakies on Hacker News, after a workday with itokay now that i’ve spent a workday with it… yeah not amazing as everyone says. It’s waaay slower than 4.5, used a lot more tokens… I had to switch to Opus to explain something to me after telling grok to do something a bunch of times and not seeing the result.
Another user remembered July precisely: “they also widely publicised their performance for 4.5 while downplaying the fact the benchmarks were ‘accidentally’ in their training set”. His verdict on daily use: a downgrade from Opus, “but barely noticeable”. A third, flagging his own guess as speculation: “around Opus 4.8 in real world use, but clearly below Opus 5.” The counterpoint exists: “Grok is 3x+ faster than Claude and I can’t tell the diff in engineering work quality.”
Both readings can be true. Artificial Analysis measured Grok 4.6 finishing agentic tasks in about 53 turns and half a billion input tokens where Opus 5 takes about 103 turns and two billion. That’s their claim, not xAI’s, and it’s the honest version of what July’s token chart was gesturing at: this model is cheap to run for long, not smartest per turn.
What This Isn’t
- Proof of reform. One clean launch after getting caught is minimum viable honesty. Unbranded columns and a vanished benchmark are still spin, just quieter.
- A flop. The 4.5-to-4.6 jumps are large: Terminal-Bench 15.7 to 26, DeepSWE 54 to 65.9. The model improved; only the story got humbler.
- A reliable calendar. Musk said on X 4.6 would land “around August 7” with the 2.1-trillion-parameter Grok 4.7 “a few weeks later”. 4.6 arrived August 12. Price the 4.7 date accordingly.
The Retake
The system worked, is the odd conclusion. Cursor caught the contamination, pulled the scores, and the next launch shipped with the full table, the clean benchmark, and the losses in plain view. Getting caught turned out to be the most effective eval-integrity intervention of the year.
But read the residue. The exam is honest now. The certificate still won’t say who came first.



