Developer
xAI

Grok 4.6

xAI's agent-focused flagship. 500K context at $2.00 in / $6.00 out per 1M — but a prompt of 200K tokens doubles the rate on the entire request.

Updated

Released
Aug 12, 2026
Type
Multimodal Language Model
Context window
500K tokens
Pricing
$2.00 / $6.00 per Mtok
License
Proprietary
On this page

Overview

Grok 4.6 is xAI's flagship model, released August 12, 2026, thirty-five days after Grok 4.5. It is a further training run on the 4.5 line rather than a new model family: the knowledge cutoff did not move, the context window did not move, and the headline price did not move. What moved is agentic performance.

The framing xAI leads with is long-running agents — work that spans many steps, accumulates context, and has to survive its own mistakes. The vendor's second theme is interactive and visual output, where xAI claims Grok 4.6 produces stronger first passes than 4.5 on things like working front-ends. Both themes are consistent with the benchmark set it chose to publish, which is dominated by agent harnesses rather than single-turn reasoning.

The Cursor relationship survived the version bump and the corporate one. Cursor says it trained Grok 4.6 jointly with SpaceXAI, as it did 4.5, and SpaceX closed its $60 billion all-stock acquisition of Anysphere — Cursor's maker — in mid-August 2026, days after this model shipped. The developer-tool funnel and the model are now the same company. Read the CursorBench numbers with that in mind.

The thing that will surprise you is the bill. The 500K window is real, but the price is not flat across it. The moment a prompt reaches 200,000 tokens, the entire request — every input, cached and output token in it — bills at double the headline rate. On a long agent loop this is the dominant cost fact about the model, and it is not visible in the $2.00 / $6.00 figure everyone quotes.

A note on the vendor. Following the SpaceX acquisition, x.ai presents the company as SpaceXAI, and the developer documentation has followed. The API endpoint is unchanged at api.x.ai. This page lists the developer as xAI, the name Grok shipped under.

Capabilities

  • Long-horizon agent loops: the stated design target. Artificial Analysis measures Grok 4.6 resolving its agentic tasks in roughly 53 turns and 0.5B input tokens against ~103 turns and ~2.0B for Claude Opus 5 at max effort — turn efficiency that compounds into cost.
  • Agentic tool calling: function calling, web search, X search and code execution are all documented capabilities of the model itself.
  • Configurable reasoning effort: low, medium, high (default) and a new xhigh level that 4.5 did not have.
  • Coding: the benchmark suite xAI leads with, and where the Cursor co-training shows — CursorBench v3.2 rises to 69.9% from 66.7%.
  • Interactive and visual work: xAI claims stronger first drafts on visual and interactive projects than 4.5, and better self-testing before it hands work back.
  • Vision input: text and image in, text out, per xAI's own model page — a documented capability this time, where 4.5's was inferred from third parties.

Technical Specifications

  • Model ID: grok-4.6
  • Architecture: Mixture of experts, per the Grok 4.5 launch materials for the line xAI continued training here. xAI has not restated the architecture for 4.6.
  • Context window: 500,000 tokens
  • Max output: no separate text output cap is published; output is bounded by the context window
  • Input modalities: Text, image
  • Output modalities: Text
  • Knowledge cutoff: February 1, 2026 per xAI's models table (the standalone model page says January 2026 — the vendor pages disagree)
  • Reasoning effort: low, medium, high (default), xhigh
  • Capabilities: Function calling, web search, X search, code execution, reasoning
  • APIs: Responses API and Chat Completions
  • API base URL: https://api.x.ai/v1

There is no Fast, Mini or Heavy variant of Grok 4.6. The "double price for a faster model" line in launch coverage refers to Priority Processing — service_tier: "priority" on the same model ID, a flat 2× multiplier on every token type, billed only when the response comes back confirming the priority tier.

xAI's docs recommend setting a prompt_cache_key to get reliable cache hits rather than paying full input price on a cache-cold server. On a model whose selling point is long agent loops, that is not a micro-optimisation.

Use Cases

  • Long-running coding agents: the model's strongest published axis and the reason Cursor ships it across desktop, web, iOS, CLI and its SDK.
  • Multi-step research and analysis agents: the turn-efficiency result matters most where each turn drags a growing transcript behind it.
  • Front-end and interactive prototyping: xAI's claimed improvement over 4.5 is specifically on first-pass visual and interactive output.
  • Cost-sensitive frontier work: at launch Artificial Analysis put it at frontier intelligence for a fraction of the cost per task of Claude Opus 5 or GPT-5.6 Sol.
  • Anything with a prompt near 200K tokens: the one case to route away. Either compact the context below the threshold or price the request at double.
  • Whole-corpus prompts: still Grok 4.3's job, at 1M tokens.

Performance / Benchmarks

Vendor-published

From xAI's own launch materials, at high reasoning effort. These are self-reported and not independently replicated.

BenchmarkGrok 4.6Grok 4.5
CursorBench v3.269.9%66.7%
DeepSWE v1.165.9%54.0%
FrontierCode v1.1 (Extended)61.3%56.6%
APEX-Agents57.5%47.1%
APEX-SWE56.4%53.6%
Terminal-Bench v3.026.0%15.7%
Harvey LAB (Vals)15.8%12.9%

The two double-digit jumps — DeepSWE v1.1 and APEX-Agents — are both agent-harness evaluations, which is the release's whole argument. The single-turn gains are ordinary. Terminal-Bench v3.0 at 26.0% is worth noticing for the opposite reason: on the harder, newer version of that benchmark, the frontier is still bad at this.

Independent, from Artificial Analysis

Measured on August 12, 2026 under Intelligence Index v4.1:

Metric (third-party, v4.1)Grok 4.6Grok 4.5
Intelligence Index6156
GDPval-AA v2 (Elo)17531526
AA-Briefcase (Elo)15771313
Terminal-Bench v2.188.4%—
τ³-Banking50.7%—

At that point 61 put Grok 4.6 level with GPT-5.6 Sol, one point behind Claude Fable 5 at 62 and two behind Claude Opus 5 at 63 — the cheapest model at the intelligence frontier, at roughly $0.84 cost per task where the Anthropic and OpenAI flagships cost multiples of that.

Then the ruler changed. Artificial Analysis published Intelligence Index v4.2 on September 4, 2026 — it adds AA-Briefcase and a 4,592-page long-context document evaluation, drops the saturated GPQA Diamond, and raises the weight of private held-out sets from 20% to 40%. Under v4.2, Grok 4.6 (high) is listed at 51 and Grok 4.5 at 45. The generational gap survives; the absolute numbers do not transfer. A v4.1 score and a v4.2 score are not comparable, and a "61" quoted today without its version is not a fact about this model.

Two frontier releases have also landed since the launch comparison above: OpenAI's GPT-6 Astra on September 4 and Anthropic's Claude Fable 5.1 on September 1. The August ranking predates both.

Grok 4.5 and Grok 4.3 Are Still Here

xAI has announced no retirement for either, and both remain in the model and pricing tables.

Grok 4.6Grok 4.5Grok 4.3
ReleasedAug 12, 2026Jul 8, 2026Apr 30, 2026
Context window500K500K1M
Input / output per 1M$2.00 / $6.00$2.00 / $6.00$1.25 / $2.50
Cached input per 1M$0.50$0.30$0.20
Intelligence Index (v4.2)5145—

There is almost no reason to stay on 4.5: same window, same headline price, worse scores. The one exception is a cache-dominated workload, where 4.5's $0.30 cached input undercuts 4.6's $0.50. Grok 4.3 remains the choice for genuinely long context — 1M tokens at $1.25 input is a combination the 4.x flagships do not offer at any price, and it does not have the 200K billing cliff.

Limitations

  • The 200K billing cliff. Reaching 200,000 prompt tokens doubles the rate on the entire request, cached tokens included. The usable-at-headline-price window is 200K, not 500K.
  • Slow. Artificial Analysis measures 63.6 tokens/sec output and a time to first token above 52 seconds — high even among reasoning models in this price tier.
  • Benchmark scores are self-reported. Every figure in the vendor table above comes from xAI, at effort settings xAI chose. No independent replication of DeepSWE, APEX or CursorBench results exists.
  • CursorBench is not a neutral benchmark. It is published by the company that co-trained the model and is now owned by the model's parent.
  • Conflicting knowledge cutoff. Two xAI pages give two dates a month apart, and neither is later than February 2026.
  • No architecture disclosure for 4.6. The mixture-of-experts description comes from the 4.5 materials.
  • Terminal-Bench v3.0 at 26.0%. A reminder that "frontier agentic model" does not mean reliable on hard multi-step terminal work.
  • Text output only. Image and video generation stay in the separate grok-imagine line; voice in grok-voice.
  • Regulatory exposure. UK and EU regulators have open investigations into the Grok consumer product; these do not currently restrict API access, but they are live.

Pricing & Access

API pricing (per 1M tokens)

Prompt under 200KPrompt at or above 200K
Input$2.00$4.00
Cached input$0.50$1.00
Output$6.00$12.00

The threshold applies to the whole request, not the overage. A prompt of 200,001 tokens does not bill 200K tokens at $2.00 and one token at $4.00 — it bills all 200,001 at $4.00, and every output and cached token in that request at the long-context rate too. Crossing the line by a single token doubles the invoice for the call. On an agent loop that grows its own context, budget for the cliff or compact before you hit it.

Priority Processing is a separate 2× multiplier on top of whichever tier applies, requested with service_tier: "priority" and charged only when the response confirms it. Cache discounts apply before the multiplier.

Access

  • API: https://api.x.ai/v1 — Responses API and Chat Completions
  • Grok Build: xAI's developer surface, included in the $30/month SuperGrok plan
  • Cursor: available across desktop, web, iOS, CLI and SDK from launch day, as co-developer
  • Gateways: OpenRouter, Vercel and Cloudflare from day one

Model Retirements & Migration

The May 15, 2026 retirement is still the only migration event xAI has run on this line, and it targets the Grok 4.x generation that preceded 4.3. On that date xAI retired eight models:

Retired modelMigration target
grok-4-1-fast-reasoninggrok-4.3 (low reasoning effort)
grok-4-1-fast-non-reasoninggrok-4.3 (none reasoning effort)
grok-4-fast-reasoninggrok-4.3 (low reasoning effort)
grok-4-fast-non-reasoninggrok-4.3 (none reasoning effort)
grok-4-0709grok-4.3 (low reasoning effort)
grok-3grok-4.3 (none reasoning effort)
grok-code-fast-1grok-build-0.1
grok-imagine-image-progrok-imagine-image-quality

The pattern was consolidation: the fast/non-fast and reasoning/non-reasoning axes collapsed into a single flagship with a reasoning-effort dial, and coding work moved to a dedicated grok-build line. Grok 4.6 does not change any of these targets, and adds no retirements of its own.

Ecosystem & Tools

  • xAI developer documentation - Model reference, pricing, migration guides
  • Grok 4.6 model page - Canonical specs and the pricing thresholds
  • OpenAI-compatible API - Reachable through OpenAI-compatible SDKs by overriding base_url to https://api.x.ai/v1
  • Composer 2.5 - Cursor's own coding model, from the company that co-trained Grok 4.6
  • grok-build-0.1 - xAI's dedicated coding model, 256K context
  • grok-imagine-image-2.0 / grok-imagine-video-1.5 - Image and video generation, separate from the Grok text line
  • grok-voice-think-fast-2.0 - Voice, billed per minute of audio

Community & Resources

Frequently Asked Questions

August 12, 2026 — 35 days after Grok 4.5. Unlike the 4.5 launch, xAI's own announcement page is readable this time, so the vendor claims on this page come from the vendor rather than from press relay.
grok-4.6. There is no Fast, Mini or Heavy variant: the doubled price sometimes described as a fast tier is Priority Processing, requested with service_tier: "priority" on the same model ID.
500,000 tokens, unchanged from Grok 4.5. Note that you cannot use the top 300K of it at the headline price — billing switches tiers at a 200,000 token prompt.
$2.00 per million input tokens, $0.50 cached, and $6.00 per million output tokens for prompts under 200,000 tokens. At or above 200,000 prompt tokens the whole request bills at $4.00 / $1.00 / $12.00 — not just the tokens past the threshold.
No. grok-4.5 remains in xAI's model and pricing tables with the same 500K window and the same headline $2.00 / $6.00 rates. Its cached-input rate is actually lower than 4.6's, at $0.30 per million against $0.50.
Yes. Cursor states that it trained Grok 4.6 jointly with SpaceXAI, continuing the arrangement behind Grok 4.5. SpaceX closed its $60 billion acquisition of Anysphere, Cursor's maker, in mid-August 2026 — days after this model shipped.
61 under Intelligence Index v4.1, the index in force at launch, level with GPT-5.6 Sol and behind Claude Fable 5 at 62 and Claude Opus 5 at 63. Artificial Analysis then published v4.2 on September 4, 2026, under which Grok 4.6 is listed at 51 and Grok 4.5 at 45. The two index versions are not comparable.
xAI's models table gives February 1, 2026, the same date it gives for Grok 4.5 — 4.6 is a further training run on that line, not a new pretrain. The standalone grok-4.6 model page says January 2026; the two vendor pages disagree by a month.

Explore More Models

Discover other AI models and compare their capabilities.