GEMINI 2.5 FLASH

Gemini 2.5 Flash: price-performance generalist

1M-token context. Multimodal. Thinking mode. Low latency. Native googleSearch grounding. $0.30/$2.50 per million tokens — well-rounded for large-scale processing and agentic use cases.

Try Gemini 2.5 Flash on Council See it in a council

Gemini 2.5 Flash is Google's previous-generation fast tier and remains one of the most well-rounded price-performance models in mid-2026. It carries the same 1M-token native context window as the Pro tier, full multimodal input (text + image + audio + video), thinking-mode reasoning when needed, and native googleSearch grounding — all at $0.30 per million input and $2.50 per million output. The case for: large-scale processing pipelines, low-latency chat, agentic use cases that need tool-calling without breaking the budget, and workloads already tuned against it. The case against: Gemini 3 Flash ($0.50/$3) is a meaningful step up in reasoning and instruction-following for ~70% more cost, and Flash Lite 3.1 ($0.25/$1.50) is cheaper if you don't need the deeper reasoning. 2.5 Flash sits between them and remains a strong default for known-good workloads.

Specs

ProviderGoogle DeepMind
Context window1,000,000 tokens (native)
ReasoningThinking mode available
MultimodalText + image + audio + video
Web searchNative googleSearch grounding
Price (in / out per 1M)$0.30 / $2.50
Tool callingStrong; agent-ready

Best at

Where it loses

2.5 Flash vs 3 Flash vs Flash Lite 3.1

For new workloads, default to Gemini 3 Flash ($0.50/$3) — it's smarter across the board for ~70% more cost than 2.5 Flash. Keep 2.5 Flash ($0.30/$2.50) for workloads already validated against it. Pick Flash Lite 3.1 ($0.25/$1.50) when you need the cheapest possible frontier-grade model — budget councils, bulk classification, anywhere quality bar is moderate.

Frequently asked questions

Is Gemini 2.5 Flash a thinking model?

It supports thinking mode on demand — you can request extended chain-of-thought when needed and skip it for fast responses. The toggle is per-request.

How does 2.5 Flash compare to 3 Flash?

Gemini 3 Flash is a clear upgrade in reasoning, instruction-following, and tool use for ~70% more cost ($0.50/$3 vs $0.30/$2.50). For most new work, default to 3 Flash.

How much does Gemini 2.5 Flash cost?

$0.30 per million input tokens and $2.50 per million output tokens. Council AI bundles it inside every paid plan; no separate Google billing.

Is 2.5 Flash good for agents?

Yes — strong tool-calling and low latency make it agent-ready at a budget price. For mission-critical agent loops, GPT-5.5 remains the most reliable, but 2.5 Flash is solid for everyday automations.

2.5 Flash or Flash Lite 3.1?

Pick 2.5 Flash when you need thinking mode or deeper reasoning. Pick Flash Lite 3.1 for raw throughput and free-tier traffic where the cost difference matters more than the reasoning gap.