In July I wrote that Grok 4.5 trained on the answer key: an earlier snapshot of the Cursor codebase, the thing CursorBench grades against, turned up in the training data, and Cursor pulled the scores. Five weeks later, Grok 4.6’s launch page does something I didn’t expect. It publishes a ten-row benchmark table that its rival wins.

Count them. The Fable 5 Max column takes six of ten rows, including the Artificial Analysis Intelligence Index (62 to Grok’s 61, with Opus 5 at 63 above both). DeepSWE, which Grok loses to GPT-5.6 Sol by 7.1 points, is on the hero chart. And CursorBench is back: reinstated after the contamination fix, scored clean at 69.9%, and losing to Fable’s 70.5%. The benchmark that got them caught, retaken honestly, and lost.

The wins are real too: the GDPVal-AA knowledge-work eval, AA-Briefcase, and a genuinely strange 15.8% on the Harvey legal benchmark against Sol’s 2.5%. At $2 in and $6 out per million tokens, the price argument from July still stands.

The Benchmark That Vanished

Forced honesty has a residue, and it’s what the page leaves out. SWE-bench Pro, the stage for July’s token-efficiency misdirection, is absent from the page entirely. Every other benchmark from the 4.5 story came back, including the one that caused the scandal. The one that anchored the misleading hero chart did not.

Quote Terminal-Bench with its version number

xAI reports Terminal-Bench v3.0 at 26%. Artificial Analysis reports Terminal-Bench v2.1 at 88.4%. Same benchmark family, different versions, wildly different scales. Any comparison that drops the version number is accidentally lying.

The Street Price

The table says tied with Sol. The first-hand reports on Hacker News, a day or two into real use, say something else.

okay now that i’ve spent a workday with it… yeah not amazing as everyone says. It’s waaay slower than 4.5, used a lot more tokens… I had to switch to Opus to explain something to me after telling grok to do something a bunch of times and not seeing the result.

— bakies on Hacker News, after a workday with it

Another user remembered July precisely: “they also widely publicised their performance for 4.5 while downplaying the fact the benchmarks were ‘accidentally’ in their training set”. His verdict on daily use: a downgrade from Opus, “but barely noticeable”. A third, flagging his own guess as speculation: “around Opus 4.8 in real world use, but clearly below Opus 5.” The counterpoint exists: “Grok is 3x+ faster than Claude and I can’t tell the diff in engineering work quality.”

Both readings can be true. Artificial Analysis measured Grok 4.6 finishing agentic tasks in about 53 turns and half a billion input tokens where Opus 5 takes about 103 turns and two billion. That’s their claim, not xAI’s, and it’s the honest version of what July’s token chart was gesturing at: this model is cheap to run for long, not smartest per turn.

What This Isn’t

  • Proof of reform. One clean launch after getting caught is minimum viable honesty. A vanished benchmark is still spin, just quieter.
  • A flop. The 4.5-to-4.6 jumps are large: Terminal-Bench 15.7 to 26, DeepSWE 54 to 65.9. The model improved; only the story got humbler.
  • A reliable calendar. Musk said on X 4.6 would land “around August 7” with the 2.1-trillion-parameter Grok 4.7 “a few weeks later”. 4.6 arrived August 12. Price the 4.7 date accordingly.

The Retake

The system worked, is the odd conclusion. Cursor caught the contamination, pulled the scores, and the next launch shipped with the full table, the clean benchmark, and the losses in plain view. Getting caught turned out to be the most effective eval-integrity intervention of the year.

But read the residue. The exam is honest now. One question from July still isn’t on the paper.