GM
Google: Gemini 2.5 Flash
Textgemini-2.5-flash
Fast, low-cost Gemini for high-throughput workloads, 1M context.
Throughput102 tok/s
24hnow
Latency0.28 s
24hnow
Specifications
Authorgoogle
ModalityText
Context1M tokens
Input$0.27/M tokens$0.30/M tokens
Output$2.25/M tokens$2.50/M tokens
Savings−10%
Capabilitiestool calling, image input
Quick start
Same OpenAI-compatible endpoint — just swap base_url and key.
API
OpenAI-compatible chat completions endpoint, with SSE streaming.
POST
/v1/chat/completionsBody parameters
Structured fields (tools, response_format, stop and more) are in the full parameter reference.
Response
Example
Errors
More from Google
Gemini 3.1 Flash Lite
gemini-3.1-flash-lite
Gemini 3 Flash
gemini-3-flash
Gemini 3.5 Flash
gemini-3.5-flash
Gemini 2.5 Pro
gemini-2.5-pro
Gemini 3.1 Pro
gemini-3.1-pro
Gemini 3 Pro
gemini-3-pro
Nano Banana Fast
nano-banana-fast
Nano Banana
nano-banana
Nano Banana 2
nano-banana-2
Nano Banana Pro
nano-banana-pro
Veo 3.1
veo-3.1
Veo 3.1 Fast
veo-3.1-fast