GM
Google: Gemini 3.1 Flash Lite
Textgemini-3.1-flash-lite
The cheapest, fastest Gemini for high-volume tasks, 1M context.
Throughput161 tok/s
24hnow
Latency1.18 s
24hnow
Specifications
Authorgoogle
ModalityText
Context1M tokens
Input$0.22/M tokens$0.25/M tokens
Output$1.35/M tokens$1.50/M tokens
Savings−10%
Capabilitiestool calling, image input
Quick start
Same OpenAI-compatible endpoint — just swap base_url and key.
API
OpenAI-compatible chat completions endpoint, with SSE streaming.
POST
/v1/chat/completionsBody parameters
Structured fields (tools, response_format, stop and more) are in the full parameter reference.
Response
Example
Errors
More from Google
Gemini 2.5 Flash
gemini-2.5-flash
Gemini 3 Flash
gemini-3-flash
Gemini 3.5 Flash
gemini-3.5-flash
Gemini 2.5 Pro
gemini-2.5-pro
Gemini 3.1 Pro
gemini-3.1-pro
Gemini 3 Pro
gemini-3-pro
Nano Banana Fast
nano-banana-fast
Nano Banana
nano-banana
Nano Banana 2
nano-banana-2
Nano Banana Pro
nano-banana-pro
Veo 3.1
veo-3.1
Veo 3.1 Fast
veo-3.1-fast