All models
gcp_vertex

Gemini 3.1 Pro

Our cloud
Try in Playground

Gemini 3.1 Pro is built to refine the performance and reliability of the Gemini 3 Pro series, with better thinking, improved token efficiency, and a more grounded, factually consistent experience. It is optimized for software engineering behavior and for agentic workflows that require precise tool usage and reliable multi-step execution.

The real Google model

Same weights, same API, same answers. We change only how you pay for it: one key for every maker, one bill, no separate account to open. Every figure on this page comes from Google’s own page — check it.

Maker's own page

Price

per 1M tokens

Input

Output

$2.00

$12.00

Long prompts

$4.00 / $18.00

Above 200,000 input tokens Google charges more, for the whole request including its reply. We pass that through unchanged.

Context

1,048,576

up to 65,536 out

Modalities

ToolsReasoning

Released

Feb 19, 2026

What it would cost you

1000 tokens is roughly 750 English words.

Estimated total$10.00

Where this model runs

gcp_vertex
Google Vertex
google
Google

If one is unavailable the request goes to the next one, and you never see it happen.

Similar models

Models built for the same kind of work, at a similar price.

We do not mark up the price of tokens. It is the same price the provider charges. We earn on the top-up fee.

Use it from your code

Model:gemini-3.1-pro
main.py
OpenAI Specification v1.0
from openai import OpenAI

# Works with the official openai package (pip install openai)
client = OpenAI(
    base_url="https://api.flintbeam.com/v1",
    api_key="sk-live-your-api-key"
)

response = client.chat.completions.create(
    model="gemini-3.1-pro",
    messages=[
        {"role": "system", "content": "You are an experienced software engineer."},
        {"role": "user", "content": "Describe the architecture of a distributed cache."}
    ],
    temperature=0.7,
    max_tokens=1500
)

print(response.choices[0].message.content)