Gemini 3.5 Flash

GoogleTextReasoningToolsVision

The newest fast Gemini 3.5 - strong reasoning at low latency.

Our price
$1.35 / $8.10
input / output per 1M
Official
$1.50 / $9.00
−10% vs official

First request

from openai import OpenAI

client = OpenAI(
    base_url="https://api.altrouter.ai/v1",
    api_key="ar-...",
)

resp = client.chat.completions.create(
    model="gemini-3.5-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

The interface is OpenAI-compatible: in existing code only the base_url, the key and the model id change — nothing else.

Gemini 3.5 Flash against its price neighbours

Gemini 3.5 Flash and its closest neighbours by price: prices and parameters from the AltRouter AI catalog
ModelInput / 1MOutput / 1MContext
GPT-5.6 MiniOpenAI$0.85$5.10400KOpen
Grok 4.6xAI$1.70$5.10256KOpen
Gemini 3.5 Flashthis pageGoogle$1.35$8.101M
Claude Sonnet 5Anthropic$1.69$8.50200KOpen
Gemini 3.1 ProGoogle$1.46$10.201MOpen

Prices and parameters come from the same catalog as the price above; this is not a benchmark, it is comparable billing terms.

Limits

Context is 1M tokens per request, and it holds everything you send: the system prompt, the conversation history and the current question together. Whatever does not fit is not seen by the model — trimming the history is your side of the job, and you pay for exactly what fit.

Response length is set by max_tokens, capped at 32,000 tokens per call. A response that hits the cap comes back with finish_reason: "length" — truncated but still billed, so it is worth setting deliberately on long generations.

Supported: a reasoning mode, tool calling, image input.

There is no subscription: every successful request is billed at the price above, and a failed one is not billed at all. That price is already 10% below the vendor's official one — the official price is shown struck through next to it so it can be checked. No Google account of your own is needed: requests go to api.altrouter.ai and we make the upstream call.

Rate limiting is shared across models and counted per key: 600 requests per minute by default, and a 429 with retry-after: 60 above it. An individual key can get its own limit and a spending cap per day, week, month or lifetime in the dashboard.

Full limits and credits

When to pick Gemini 3.5 Flash, and when a neighbour

The cheaper neighbour is Grok 4.6: $1.70 per 1M input tokens against $1.35 for Gemini 3.5 Flash. Its context is smaller: 256K against 1M — long documents will need splitting. It is the sensible pick where the task is repetitive and high-volume — labelling, classification, short templated answers — the kind of work where quality is bounded by the prompt rather than the model.

The dearer neighbour is Claude Sonnet 5: $1.69 per 1M input tokens, or 1.3× the price. On catalog parameters it offers the same capabilities and the same order of context, so the premium is only justified if Gemini 3.5 Flash falls short on your own task.

Guessing is more expensive than checking: the table above is prices and parameters, not a benchmark, and it cannot tell you which model does better on your own prompt. Switching costs one line — the model id in the request changes, the base_url and the key stay — so running both on your own data takes minutes.

API

OpenAI-compatible chat completions endpoint, with SSE streaming.

POST/v1/chat/completions
Body parameters
modelstringrequired"gemini-3.5-flash"
The model id.
messagesarrayrequired
The conversation: [{ role, content }].
streambooleandefault false
Stream the response over SSE.
temperaturenumber0–2default 1
Sampling randomness: lower is more deterministic, higher more creative.
top_pnumber0–1default 1
Nucleus sampling - the probability mass to sample from.
max_tokensnumber1–32000
Maximum tokens to generate.
reasoning_effortenumlow · medium · highdefault medium
Thinking budget for reasoning models.
seednumber0–2147483647
Seed for reproducibility.
frequency_penaltynumber-2–2
Penalizes tokens by their existing frequency.
presence_penaltynumber-2–2
Penalizes tokens that have already appeared.
web_searchbooleandefault false
Server-side web search - answers with live data.

Structured fields (tools, response_format, stop and more) are in the full parameter reference.

Response
idstring
Completion id.
choicesarray
Completions. Each: index, message { role, content, reasoning_content, tool_calls }, finish_reason.
usageobject
Tokens: prompt_tokens, completion_tokens, total_tokens.
Example
cURLPythonJavaScriptJSON
curl https://api.altrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer ar-..." \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gemini-3.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Hello!"
    }
  ],
  "reasoning_effort": "medium"
}'
Errors
400Invalid request.
401Missing or invalid API key.
402Insufficient credits, or the key’s spending cap is exhausted.
404Model not found.
429Rate limit exceeded.

More from Google