GEMINI 2.5 FLASH
1M-token context. Multimodal. Thinking mode. Low latency. Native googleSearch grounding. $0.30/$2.50 per million tokens — well-rounded for large-scale processing and agentic use cases.
Gemini 2.5 Flash is Google's previous-generation fast tier and remains one of the most well-rounded price-performance models in mid-2026. It carries the same 1M-token native context window as the Pro tier, full multimodal input (text + image + audio + video), thinking-mode reasoning when needed, and native googleSearch grounding — all at $0.30 per million input and $2.50 per million output. The case for: large-scale processing pipelines, low-latency chat, agentic use cases that need tool-calling without breaking the budget, and workloads already tuned against it. The case against: Gemini 3 Flash ($0.50/$3) is a meaningful step up in reasoning and instruction-following for ~70% more cost, and Flash Lite 3.1 ($0.25/$1.50) is cheaper if you don't need the deeper reasoning. 2.5 Flash sits between them and remains a strong default for known-good workloads.
For new workloads, default to Gemini 3 Flash ($0.50/$3) — it's smarter across the board for ~70% more cost than 2.5 Flash. Keep 2.5 Flash ($0.30/$2.50) for workloads already validated against it. Pick Flash Lite 3.1 ($0.25/$1.50) when you need the cheapest possible frontier-grade model — budget councils, bulk classification, anywhere quality bar is moderate.
It supports thinking mode on demand — you can request extended chain-of-thought when needed and skip it for fast responses. The toggle is per-request.
Gemini 3 Flash is a clear upgrade in reasoning, instruction-following, and tool use for ~70% more cost ($0.50/$3 vs $0.30/$2.50). For most new work, default to 3 Flash.
$0.30 per million input tokens and $2.50 per million output tokens. Council AI bundles it inside every paid plan; no separate Google billing.
Yes — strong tool-calling and low latency make it agent-ready at a budget price. For mission-critical agent loops, GPT-5.5 remains the most reliable, but 2.5 Flash is solid for everyday automations.
Pick 2.5 Flash when you need thinking mode or deeper reasoning. Pick Flash Lite 3.1 for raw throughput and free-tier traffic where the cost difference matters more than the reasoning gap.