Grok 4.6 posts a 1753 and a number one. One is a real benchmark. The other is a tweet.
The 1753 is genuine and Grok is genuinely cheap. But it is tied for third overall, and the number one Musk tweeted is missing from both xAI's published table and Databricks' own leaderboard.
"Grok 4.6 reaches 1753 ELO and is #1 on the Databricks leaderboard, at half the price of rival frontier models." [SOURCE ↗]

THE CLAIM. Musk says Grok 4.6 reaches 1753 ELO and is number one on Databricks' OfficeQA benchmark, at half the price of rivals.
THE CHECK. the 1753 is real, and it comes from Artificial Analysis, not just Musk. But it is one benchmark, GDPval-AA v2, where Grok sits behind Claude Opus 5 and inside overlapping confidence intervals with Fable 5 and Qwen. On the overall Intelligence Index it scores 61, tied with GPT-5.6 Sol for third, behind Opus at 63 and Fable at 62.
THE TWIST. the number one Musk tweeted is an OfficeQA Pro V2 run with Databricks' Genie harness, and in one analysis's words that number 'is not in SpaceXAI's published table.' xAI's model card does show a number one on the older OfficeQA Pro v1, 63.2 percent, but xAI ran that itself, and Databricks' own leaderboard does not list Grok 4.6 at all. The part that actually holds is the boring part: Grok 4.6 is cheap, about $0.84 a task, with list prices of $2 in and $6 out per million tokens against $5 and $25 for Opus 5.
What actually happened
SpaceXAI, formerly xAI, released Grok 4.6, and Elon Musk posted that it 'reaches 1753 ELO' and is number one on the Databricks leaderboard, at half the price of rival frontier models. The launch was the day's biggest, and unusually for a Musk number, the headline figure is real.
Artificial Analysis, an independent evaluator, confirms it: Grok 4.6 'achieves a GDPval-AA v2 Elo of 1753.' That is a genuine result from a third party, not a self-report. So far, so good.
Why we rate this needs_context
Three things the tweet leaves out.
First, 1753 is one benchmark. GDPval-AA v2 measures real-world agentic tasks, and on it Grok sits 'behind only Claude Opus 5,' with confidence intervals that overlap Claude Fable 5 and Qwen3.8 Max. Overlapping intervals means statistically tied, not beaten.
Second, the overall picture is a tie for third, not first. On Artificial Analysis's headline Intelligence Index, Grok 4.6 'scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62).' A tie for third, by a nose, is a fine result. It is not a number one.
A single benchmark is a spotlight, and you can always stand where the spotlight is.
Third, and sharpest: the number one Musk tweeted is a specific run, OfficeQA Pro V2 with Databricks' Genie harness, and that result has no paper trail. As one launch analysis puts it flatly, 'That number is not in SpaceXAI's published table.' The advice from the same piece: 'Treat it as a founder claim until Databricks or a third party reproduces it.' To be precise about what does exist: xAI's own model card (the Grok 4.6 PDF on media.x.ai) shows a number one on the older OfficeQA Pro v1, 63.2 percent against 60.9 for Claude Opus 5, but xAI ran that benchmark itself. Databricks' official OfficeQA leaderboard, last updated August 2, does not list Grok 4.6 at all.
The part that actually holds
Strip the ranking theater and there is a genuine win underneath: cost. Grok 4.6 runs at roughly $0.84 per task on Artificial Analysis's harness, the same as Kimi K3, and its list prices, $2 in and $6 out per million tokens, sit well below Opus 5 at $5 and $25 and GPT-5.6 Sol at $5 and $30. If your decision is price-per-capability, Grok 4.6 is a real option, and that is the claim xAI could have led with honestly.
The steelman, and why the framing still needs the caveat
The fair case: the jump from Grok 4.5 to 4.6 is large and real, the 1753 is independently confirmed, and being third in the frontier cluster while undercutting everyone on price is a legitimately strong position. Agreed. The problem is only the compression: 'third, tied, on one benchmark, plus an unverified Databricks crown' becomes '1753, number one.' The numbers are real; the ranking is chosen.
The mechanism
Cross-model comparisons in xAI's table use the best of each rival's self-reported or publicly available results, not one lab running all four models in the same harness with the same tools and settings. That means the table is assembled from different conditions, which is exactly how a third-place model ends up sounding like a leader. The founder tweet does the rest.
What to do with this
- Separate the three claims: 1753 (real, one benchmark), number-one-Databricks (unverified tweet), and half-price (real on list prices). Believe the first and third, hold the second.
- Distrust any leaderboard result assembled from each model's 'best publicly available score.' Wait for one evaluator to run them all under identical conditions.
- If you are choosing on cost, Grok 4.6 is worth testing on your own workload. If you are choosing on raw intelligence, the top of the cluster is still Opus and Fable.
The leaderboard game is won by choosing the leaderboard. A real number and a real strength, price, get bundled with an unverified crown so the whole thing reads as frontier dominance. Buy the cost efficiency, not the coronation, and wait for one lab to run every model in one identical harness before you believe any ranking.
When an independent evaluator runs Grok 4.6 and its rivals in one identical harness, Grok lands in the frontier cluster on price and mid-pack on raw intelligence, not alone at number one. The cost win survives, the crown does not. Hold us to it.
Flips to a real lead if Databricks or an independent party reproduces the OfficeQA Pro V2 number-one result, or if Grok 4.6 tops a major public benchmark run under identical conditions for every model.
RECEIPTS (9) · CONFIDENCE HIGH
every URL below answered a live HTTP check before publish · sweep 2026-08-28
- ● artificialanalysis.ai ⧉ · "achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5"
- ● artificialanalysis.ai ⧉ · "It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62)"
- ▼ explainx.ai ⧉ · "That number is not in SpaceXAI's published table."
- ● explainx.ai ⧉ · "Treat it as a founder claim until Databricks or a third party reproduces it."
- ▲ cryptobriefing.com ⧉ · "Grok 4.6 scored 1,753 on GDPVal AA v2 compared with 1,728 for GPT 5.6 Sol Max"
- ● github.com ⧉ · "Opus 5 and Antigravity Gemini 3.1 Pro (denoted by *) results updated August 2 2026"
- ● github.com ⧉ · "Headline results on OfficeQA Pro (N=133), followed by OfficeQA Pro V2 (N=90)"
- ▲ eesel.ai ⧉ · "charging $2/$6 per million tokens against Sol's $5/$30"
- ▲ eesel.ai ⧉ · "Claude Opus 5 is $5/$25"









