Introduction
xAI released Grok 4.6 on August 12, 2026. Three weeks later, two things about that release are clear that were not clear on launch day, and neither of them is a benchmark result.
The first is that Grok 4.6's headline intelligence score has fallen by ten points without a single weight changing. The second is that the model Cursor shipped in August was not a Cursor model — and two days after it shipped, Cursor stopped being an independent company. Both stories are about attribution: what a number is attached to, and whose name a model carries. This is an analysis of what the release looks like from September, not a write-up of the announcement.
The model itself is documented on our Grok 4.6 model page. This post is about the three things around it that a launch-day post could not have said.
The Score Moved Ten Points and the Model Stayed Still
On launch day, Artificial Analysis measured Grok 4.6 at 61 on Intelligence Index v4.1 — level with GPT-5.6 Sol, one point behind Claude Fable 5 and two behind Claude Opus 5, at roughly $0.84 per task where the Anthropic and OpenAI flagships cost several times that. That combination, frontier intelligence at the bottom of the price band, was the entire commercial argument for the model, and it was the number every piece of launch coverage led with.
On September 4, 2026, Artificial Analysis published Intelligence Index v4.2. Under v4.2, Grok 4.6 is listed at 51, and Grok 4.5 at 45.
Nothing shipped in between. grok-4.6 returns the same weights it returned in August. What changed is the ruler:
- Two evaluations were added: AA-Briefcase, a private agentic knowledge-work set, and GDP.pdf, a long-context document-reasoning task spanning 4,592 PDF pages.
- One was retired: GPQA Diamond, dropped as saturated — frontier models now score high enough on it consistently that it no longer separates the field.
- The private held-out weighting doubled, from roughly 20% of the index to 40%. Artificial Analysis's stated reason is resistance to gaming: the more of the index a lab cannot train against, the more the score means.
Every one of those changes lowers scores across the board, and lowers them unevenly. That is the point. Retiring a saturated eval removes free points that every frontier model was collecting; doubling the held-out weight moves the index toward the tasks no lab has optimised for. A model that looked frontier-level under v4.1 does not automatically look frontier-level under v4.2, and the gap between two models can widen or narrow purely as an artefact of recomposition.
What survives the version change is the relative result. Grok 4.6 beat Grok 4.5 by five points under v4.1 (61 to 56) and by six under v4.2 (51 to 45). The generational improvement is real under both rulers. The absolute number is real under neither, on its own.
Why we make this a build error
This site's content validator fails the build on an Artificial Analysis Intelligence Index score written without its index version. That rule exists because of exactly this situation, and Grok 4.6 is the cleanest illustration of it in the current catalog: "Grok 4.6 scores 61" and "Grok 4.6 scores 51" are both true, both sourced from Artificial Analysis, and three weeks apart. A reader handed either one without a version has been given a number they cannot use.
The general form of the rule is worth more than the specific one. A benchmark score is a measurement made by an instrument, and the instrument has a version. When the instrument is revised — and every serious eval suite is revised, because saturation is the normal end state of a good benchmark — historical scores do not migrate. They expire. The honest citation carries the date, the version, and the harness settings, or it carries nothing.
For reference, under v4.2 the top of the board is Claude Fable 5.1 at 57 and GPT-6 Astra (max) at 55. Grok 4.6's 51 places it below both — but it is also worth remembering that both of those models shipped after Grok 4.6, on September 1 and September 4 respectively. The August comparison and the September one are not measuring the same field either.
The Pricing Cliff at 200,000 Tokens
The second thing launch coverage under-reported is where the money actually goes.
Grok 4.6's advertised price is $2.00 per million input tokens, $0.50 cached, $6.00 output. Its context window is 500,000 tokens, unchanged from Grok 4.5. Those two facts do not compose the way you would expect.
| Prompt under 200K | Prompt at or above 200K | |
|---|---|---|
| Input | $2.00 | $4.00 |
| Cached input | $0.50 | $1.00 |
| Output | $6.00 | $12.00 |
xAI's pricing table states the mechanism plainly: requests whose prompt reaches the listed token threshold are "billed at the higher rate for all tokens in the request."
Read that carefully, because it is not how tiered pricing usually works. This is not an overage rate on the tokens past 200,000. A prompt of 200,001 tokens does not bill 200,000 at $2.00 and one token at $4.00 — it bills all 200,001 at $4.00, and every output token and every cached token in that same request at the doubled rate too. One token over the line doubles the invoice for the entire call.
The practical consequence: the headline-price window of this model is 200K, not 500K. The upper 300K is available, and it costs double for the privilege of using any of it. On a long-running agent loop — the exact workload xAI positions Grok 4.6 for, and the workload that accumulates context by design — this is the dominant cost fact about the model, and it is invisible in the "$2 / $6" figure everyone quotes.
Two mitigations are worth building in from the start. Compact aggressively and keep prompts below the threshold rather than drifting across it mid-run; and set a prompt_cache_key, which xAI's own docs recommend, so that prompt caching actually hits instead of paying full input price on a cache-cold server. Cache discounts apply before any multiplier, so they compound in your favour on both sides of the cliff.
One clarification while we are on price. There is no grok-4.6-fast. xAI's model list has a single 4.6 entry. The "faster variant at twice the price" that appeared across launch coverage is Priority Processing: service_tier: "priority" on the same model ID, a flat 2x multiplier on input, cached, output and reasoning tokens, available on Chat Completions and Responses only, and billed at the priority rate only when the response comes back confirming it. Same weights, same context window, better queue position. If you have been shopping for a Fast tier, you have been shopping for a scheduling parameter.
Cursor's Newest Model Shipped Under Someone Else's Name
The third story is corporate, and it is the one with the longest shadow.
Grok 4.6 was not trained by xAI alone. Cursor's own launch page states it flatly: "We trained Grok 4.6 jointly with SpaceXAI." The model shipped inside Cursor on day one across desktop, web, iOS, CLI and the SDK, and xAI's announcement lists Cursor alongside Grok Build as a launch surface.
Then, on August 14, 2026 — two days later — Cursor published Cursor is now a part of SpaceX: "This completes the acquisition process that started in April, when we announced our partnership with SpaceXAI to accelerate our model training efforts." The same post points back at the model: "Grok 4.6, which we released Wednesday, provides an early look at what we can now build together."
Line the dates up and the sequence reads as one event, not two. The joint training arrangement announced in April produced a model in August, and the acquisition closed on the strength of it.
What that means for Cursor's own model line is the part worth sitting with. Composer 2.5 shipped in May 2026 and is still the newest model in the Composer family. At its Compile conference in June, Cursor announced "a new model — our first model trained from scratch," with no name, no version and no date; nearly three months on it appears in neither Cursor's model list, its pricing table nor its changelog. The successor that did arrive in the interval arrived wearing the Grok badge.
So Cursor now offers, in the same model picker, a house model post-trained from an open-source checkpoint it did not pretrain, and a frontier model it co-trained but does not name. Its parent company owns the frontier model. That is a coherent strategy — it is cheaper to co-train a frontier model with the lab that owns the compute than to build one — but it is a different company from the one that shipped Composer 1.
It also has an evaluation consequence that anyone quoting Grok 4.6's coding numbers should register. CursorBench v3.2, on which Grok 4.6 scores 69.9% against 4.5's 66.7%, is published by the company that co-trained the model and is now owned by the model's parent. That does not make the number false. It does mean it is not third-party evidence, and it should not be read as though it were.
What Actually Improved
Stripped of the index churn, the vendor-reported generational gains are the substance of the release. All figures below are xAI's own, at high reasoning effort, self-reported and not independently replicated:
| Benchmark | Grok 4.5 | Grok 4.6 |
|---|---|---|
| DeepSWE v1.1 | 54.0% | 65.9% |
| APEX-Agents | 47.1% | 57.5% |
| Terminal-Bench v3.0 | 15.7% | 26.0% |
| FrontierCode v1.1 | 56.6% | 61.3% |
| APEX-SWE | 53.6% | 56.4% |
| CursorBench v3.2 | 66.7% | 69.9% |
| Harvey LAB | 12.9% | 15.8% |
The shape of that table is the argument. The two largest jumps — DeepSWE at nearly twelve points and APEX-Agents at over ten — are both agent-harness evaluations, where a model has to run a long loop and survive its own mistakes. The single-turn gains are ordinary. Grok 4.6 is a fine-tune of the 4.5 line aimed squarely at long-horizon agentic work, and it hit what it aimed at.
Terminal-Bench v3.0 at 26.0% is worth noticing for the opposite reason. It is a large relative improvement over 15.7% and still an absolute score that should temper anyone's expectations about autonomous multi-step terminal work. "Frontier agentic model" and "reliable at hard terminal tasks" remain two different claims.
Beyond benchmarks, 4.6 adds an xhigh reasoning effort level that 4.5 did not have, and xAI now documents image input directly rather than leaving it to be inferred.
One Unresolved Fact: The Knowledge Cutoff
We could not reconcile this one, so we are reporting it rather than picking a side.
xAI's models and pricing table states that "The knowledge cut-off date of Grok 4.6 is February 1, 2026." The standalone grok-4.6 model page in the same documentation set states January 2026. Both are vendor pages, published by xAI, live simultaneously, a month apart. There is no third source that adjudicates, and xAI has not corrected either.
Practically the gap is small — nothing later than February 2026 is in the model either way — but it is a useful reminder that "check the primary source" assumes the primary source agrees with itself. When it does not, the honest move is to say so and cite both, not to quietly pick the one that reads better.
Conclusion
Grok 4.6 is a solid agentic fine-tune of the Grok 4.5 line: real double-digit gains on agent-harness evaluations, an xhigh effort level, documented image input, and the lowest price in its intelligence band. None of that changed in the three weeks since it shipped.
What changed is everything around it. The benchmark that made its commercial case was rewritten, and its score under the new one is ten points lower for reasons that have nothing to do with the model. Its co-developer was acquired by its parent, making the most-quoted coding benchmark in its launch table a first-party number. And the price that made it attractive turns out to apply to 40% of its context window, with a cliff rather than a slope at the boundary.
Three weeks is not a long time. It was long enough for all three.
The transferable lesson is the first one. Quote the index version or do not quote the index — and hold every other benchmark to the same standard, including the ones published by parties with an interest in the result.
Sources
- Introducing Grok 4.6 — xAI's launch announcement, August 12, 2026
- Grok 4.6 model documentation — specs, modalities, reasoning effort, and the January 2026 cutoff
- Grok models and pricing — the models table, and the February 1, 2026 cutoff
- xAI API pricing — the 200K threshold rows and the "all tokens in the request" wording
- Priority Processing —
service_tier: "priority"and the 2x multiplier - Cursor: Grok 4.6 — "We trained Grok 4.6 jointly with SpaceXAI"
- Cursor is now a part of SpaceX — August 14, 2026 acquisition announcement
- Announcing Artificial Analysis Intelligence Index v4.2 — September 4, 2026 index revision
- Artificial Analysis Intelligence Index — current v4.2 leaderboard
- Artificial Analysis: Grok 4.6 benchmarks and analysis — the launch-day v4.1 measurement