GROK 4.20

Grok 4.20: 2M-token long-context reasoning model

xAI's long-context variant. 2M-token window. Reasoning mode. Vision input. Native X / web / news search. $1.25/$2.50 per million tokens.

Try Grok 4.20 on Council See it in a council

Grok 4.20 is xAI's long-context flagship and the largest production context window in the frontier set — 2,000,000 tokens, double Gemini 3 Pro's already-massive 1M window. It pairs that window with an explicit reasoning mode for harder analytical work, vision input for diagrams and screenshots, and native X / web / news search for fresh grounding. Pricing matches Grok 4.3 at $1.25 per million input tokens and $2.50 per million output, which makes it the cheapest way to feed truly enormous prompts to a frontier model. The case for Grok 4.20: whole-monorepo analysis, multi-book research synthesis, very long video transcripts, or any workload where 1M tokens isn't enough. The case against: for everyday work the smaller Grok 4.3 is identically priced but slightly snappier, and Claude Opus 4.7 still wins on code-edit quality and writing voice. Pick Grok 4.20 when context size is the binding constraint.

Specs

ProviderxAI
Context window2,000,000 tokens
ReasoningExplicit reasoning mode
MultimodalText + image input
Web searchNative X + web + news
Price (in / out per 1M)$1.25 / $2.50
Best forHuge codebases, long-doc synthesis

Best at

Where it loses

When to pick Grok 4.20

Frequently asked questions

What's the difference between Grok 4.20 and Grok 4.3?

Grok 4.3 is the daily-driver flagship with a 1M-token window. Grok 4.20 is the long-context variant with a 2M-token window and an explicit reasoning mode. Use 4.20 when 1M tokens isn't enough or when you want the reasoning mode for hard analytical tasks.

How much does Grok 4.20 cost?

$1.25 per million input tokens and $2.50 per million output tokens — identical to Grok 4.3 despite the larger window. Council AI bundles it inside monthly plan budgets.

Can Grok 4.20 actually use the full 2M-token window coherently?

In practice, recall holds well across the window for retrieval-style questions. As with any long-context model, performance on multi-hop reasoning gradually degrades past ~1M tokens, so for very complex synthesis tasks chunk-and-merge is still worth considering.

Does Grok 4.20 support function calling?

Yes, with native tool calling. The same reliability profile as Grok 4.3 — solid for most agent flows.

Is Grok 4.20 faster or slower than Grok 4.3?

Slightly slower at large input sizes due to the longer-context routing. For short prompts the difference is negligible, but you'll feel it when feeding 500K+ tokens.

Should I pick Grok 4.20 over Gemini 3 Pro for long context?

Grok 4.20 has a larger window (2M vs 1M) and the cheaper price. Gemini 3 Pro has stronger multimodal (audio + video) and Google Search grounding. If your workload is pure text, Grok 4.20 wins on size and cost.