GEMINI 3.1 FLASH LITE

Gemini 3.1 Flash Lite: cheap, fast, multimodal

Stable Flash-Lite tier. 1M context. Native multimodal. googleSearch grounding. $0.25/$1.50 per million tokens — the cheapest frontier-grade model in the lineup.

Try Flash Lite on Council See it in a council

Gemini 3.1 Flash Lite is Google's cheapest production frontier model and a workhorse budget pick in Council AI. It keeps the 1M-token context window, the multimodal stack (text + image + audio + video), and the googleSearch grounding — but trims reasoning depth and latency to hit $0.25/$1.50 per million tokens. The 3.1 designation marks it as the stable iteration of the Flash-Lite line, with better instruction-following and tool use than the original 3.0 Flash-Lite. The case for: high-throughput tasks, classification, bulk summarization, budget councils, and anywhere you'd otherwise reach for Haiku 4.5 but want native multimodal. The case against: anything requiring deeper reasoning chains, complex coding, or refined writing — promote to Flash 3, Pro 3, or a Claude model.

Specs

ProviderGoogle DeepMind
Context window1,000,000 tokens (native)
MultimodalText + image + audio + video
Web searchNative googleSearch grounding
Price (in / out per 1M)$0.25 / $1.50
Tool callingStrong
Latency tierFlash-Lite (fastest)

Best at

Where it loses

When to use Flash Lite vs Flash vs Pro

Use Flash Lite 3.1 for any task where the cost difference matters and the quality bar is moderate — budget councils, classification, bulk summarization, simple chat. Promote to Flash 3 ($0.50/$3) when output quality starts mattering — production chat, vision QA, anything customer-facing. Promote to Pro 3 ($2.50/$10) when you need deep reasoning, not just throughput.

Frequently asked questions

Why 3.1 and not 3.0?

3.1 is the stable Flash-Lite iteration with improved instruction-following and tool use over the initial 3.0 release. Council AI defaults to 3.1 across all tiers including free.

Is Flash Lite good enough for production?

For high-volume, well-scoped tasks — yes. Classification, extraction, summarization, simple chat, bulk vision. For open-ended reasoning or coding, no — use Flash 3, Pro 3, or a Claude model.

How much does Flash Lite 3.1 cost?

$0.25 per million input tokens and $1.50 per million output tokens. Council AI bundles it inside every plan including free; no separate Google billing.

Does Flash Lite support vision and video?

Yes — same multimodal stack as the rest of the Gemini 3 line. Text, image, audio, and video inputs in a single API call.

Flash Lite or DeepSeek V4 Flash for bulk work?

DeepSeek V4 Flash is cheaper per token; Flash Lite 3.1 brings multimodal and googleSearch grounding. Pick DeepSeek for text-only bulk codegen, Flash Lite for anything involving images, audio, video, or fresh information.