Gemini 2.5 Flash

GoogleTextToolsVision

Fast, low-cost Gemini for high-throughput workloads, 1M context.

Our price
$0.27 / $2.25
input / output per 1M
Official
$0.30 / $2.50
−10% vs official

First request

from openai import OpenAI

client = OpenAI(
    base_url="https://api.altrouter.ai/v1",
    api_key="ar-...",
)

resp = client.chat.completions.create(
    model="gemini-2.5-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

The interface is OpenAI-compatible: in existing code only the base_url, the key and the model id change — nothing else.

Gemini 2.5 Flash against its price neighbours

Gemini 2.5 Flash and its closest neighbours by price: prices and parameters from the AltRouter AI catalog
ModelInput / 1MOutput / 1MContext
Gemini 3.5 Flash LiteGoogle$0.27$2.251.0MOpen
Gemini 2.5 Flashthis pageGoogle$0.27$2.251M
Claude Haiku 4.5Anthropic$0.77$4.00200KOpen
GPT-5.6 MiniOpenAI$0.85$5.10400KOpen
Grok 4.6xAI$1.70$5.10256KOpen

Prices and parameters come from the same catalog as the price above; this is not a benchmark, it is comparable billing terms.

Limits

Context is 1M tokens per request, and it holds everything you send: the system prompt, the conversation history and the current question together. Whatever does not fit is not seen by the model — trimming the history is your side of the job, and you pay for exactly what fit.

Response length is set by max_tokens, capped at 32,000 tokens per call. A response that hits the cap comes back with finish_reason: "length" — truncated but still billed, so it is worth setting deliberately on long generations.

Supported: tool calling, image input. Not supported: a reasoning mode — a request using them errors out rather than silently dropping them.

There is no subscription: every successful request is billed at the price above, and a failed one is not billed at all. That price is already 10% below the vendor's official one — the official price is shown struck through next to it so it can be checked. No Google account of your own is needed: requests go to api.altrouter.ai and we make the upstream call.

Rate limiting is shared across models and counted per key: 600 requests per minute by default, and a 429 with retry-after: 60 above it. An individual key can get its own limit and a spending cap per day, week, month or lifetime in the dashboard.

Full limits and credits

When to pick Gemini 2.5 Flash, and when a neighbour

The cheaper neighbour is Gemini 3.5 Flash Lite: $0.27 per 1M input tokens against $0.27 for Gemini 2.5 Flash. On the same volume the bill comes out 1× smaller. Its context is even larger: 1.0M against 1M. It is the sensible pick where the task is repetitive and high-volume — labelling, classification, short templated answers — the kind of work where quality is bounded by the prompt rather than the model.

The dearer neighbour is Claude Haiku 4.5: $0.77 per 1M input tokens, or 2.9× the price. What the catalog says you get for it: a reasoning mode. If your task needs none of those, the premium buys nothing.

Guessing is more expensive than checking: the table above is prices and parameters, not a benchmark, and it cannot tell you which model does better on your own prompt. Switching costs one line — the model id in the request changes, the base_url and the key stay — so running both on your own data takes minutes.

API

OpenAI-compatible chat completions endpoint, with SSE streaming.

POST/v1/chat/completions
Body parameters
modelstringrequired"gemini-2.5-flash"
The model id.
messagesarrayrequired
The conversation: [{ role, content }].
streambooleandefault false
Stream the response over SSE.
temperaturenumber0–2default 1
Sampling randomness: lower is more deterministic, higher more creative.
top_pnumber0–1default 1
Nucleus sampling - the probability mass to sample from.
max_tokensnumber1–32000
Maximum tokens to generate.
seednumber0–2147483647
Seed for reproducibility.
frequency_penaltynumber-2–2
Penalizes tokens by their existing frequency.
presence_penaltynumber-2–2
Penalizes tokens that have already appeared.
web_searchbooleandefault false
Server-side web search - answers with live data.

Structured fields (tools, response_format, stop and more) are in the full parameter reference.

Response
idstring
Completion id.
choicesarray
Completions. Each: index, message { role, content, reasoning_content, tool_calls }, finish_reason.
usageobject
Tokens: prompt_tokens, completion_tokens, total_tokens.
Example
cURLPythonJavaScriptJSON
curl https://api.altrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer ar-..." \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Hello!"
    }
  ]
}'
Errors
400Invalid request.
401Missing or invalid API key.
402Insufficient credits, or the key’s spending cap is exhausted.
404Model not found.
429Rate limit exceeded.

More from Google