Skip to content
Grok Is Making a Run at the Top of the AI Leaderboard
SpaceX
This week, the new Grok 4.6 moved from 8th to 4th place on the key AI model intelligence leaderboard. That’s noteworthy because over the last couple of years the model has been between 6th and 15th place. The move to 4th is impressive, but in the AI model leapfrog game the future position is debatable. I believe Grok’s trend in model intelligence will continue to improve relative to the competition over the next 6 months, yielding a top spot early in 2027. That’s important because in the long term, premium models will enjoy pricing leverage.

Key Takeaways

Grok 4.6 yielded Grok’s highest ranking to date.
Grok’s rapid advancement along with upcoming data infusions should land it at the top of the leaderboard early in 2027.
Reaching the top of the leaderboard matters because customers who want access to these models will have to pay a premium.
1

Grok 4.6

On SpaceX’s June earnings call, Elon highlighted that the upcoming Grok 4.6 would yield meaningful advancement in intelligence. That’s exactly what happened 8 days later when the new model went live. Within a day, the model showed its best performance on the AI leaderboard, highlighted by Artificial Analysis’s Intelligence Index, which is considered the gold standard, ranking Grok 4.6 in 4th place, behind Claude Opus 5, Claude Fable 5, and GPT-5.6 Sol.

The reason for the improvements fell into the category of training improvements and was not the result of the inclusion of Cursor or SpaceX data. The bottom line is the improvements to Grok 4.6 were a function of general model advancements, rather than the inclusion of key new data.

2

Upcoming Improvements to Grok

The recent improvements to Grok relative to Claude and GPT have yielded a shrug from many investors given the leapfrog nature of these models. The roadmap is simple: a new model comes out, leapfrogs existing models, then three weeks later a new model does the same thing, and the cycle continues. In other words, we’ve seen that to stay at the top of the leaderboard you have to innovate at a rapid pace, as in new model updates every 2-3 months.

On SpaceX’s July quarter earnings call, Elon said Grok 4.7 (as in the model after Grok 4.6) would be out in about three or four weeks (targeting roughly late August to early September) and that the cadence of AI development is expected to improve dramatically. He has since reiterated (on August 12) that Grok 4.7 is significantly better than 4.6, that initial training is complete, and that the model will include a significant increase in Cursor data, along with the continuation of preliminary SpaceX data.

Grok 5 is expected before the end of 2026 and will be “incorporating the entire corpus of SpaceX data. Basically, all the data that SpaceX has ever produced, which is a tremendous amount over the course of a quarter century.” Elon has said this should make Grok “by far the best engineer.” In other words, this model is going to be a monster. The bottom line is it’s hard to get to the top of the leaderboard and I believe Grok will accomplish this by early 2027.

3

Why The Leaderboard Matters

In the world of models leapfrogging one another, it raises the question: why does the ranking even matter? The answer is not all tokens are created equally. Looking at the current cost per task for the top 10 models underscores this dynamic. Number 10 on the most recent intelligence ranking is DeepSeek, which is $0.25/task. Compare that to #1 Claude Opus 5, which is $2.34/task (the second most expensive behind Claude Fable 5 at $3.14/task). As a point of reference, Grok 4.6 is the 6th most expensive at $0.84/task.

Note that there’s a meaningful gap between #1 Claude Opus 5 and #4 Grok 4.6, $2.34/task vs. $0.84/task, which underscores the market is willing to pay up for the highest level of intelligence. In the future, there will be token pricing wars in which model providers will discount their product to gain share. These pricing dynamics will have the effect of lowering the average token price for all models, and will result in increased usage, which will yield greater total dollars spent on tokens (the Jevons paradox). Even in this price-competitive world, the most intelligent models will be priced at a premium. The bottom line is it pays to be smart.

Disclaimer

Back To Top