Gemini 2.5 Flash-Lite
Our cloudGemini 2.5 Flash-Lite is the fastest and most cost-efficient multimodal model in the Gemini 2.5 family. Google recommends it for high-volume classification, simple data extraction, and extremely low-latency applications where budget and speed are the primary constraints. It accepts text, images, video, audio and PDF, and returns text.
The real Google model
Same weights, same API, same answers. We change only how you pay for it: one key for every maker, one bill, no separate account to open. Every figure on this page comes from Google’s own page — check it.
Maker's own pagePrice
per 1M tokens
Input
Output
$0.10
$0.40
Context
1,048,576
up to 65,536 out
Modalities
Released
Jul 2025
Knows up to Jan 2025
What it would cost you
1000 tokens is roughly 750 English words.
Where this model runs
If one is unavailable the request goes to the next one, and you never see it happen.
Similar models
Models built for the same kind of work, at a similar price.
We do not mark up the price of tokens. It is the same price the provider charges. We earn on the top-up fee.
Use it from your code
from openai import OpenAI
# Works with the official openai package (pip install openai)
client = OpenAI(
base_url="https://api.flintbeam.com/v1",
api_key="sk-live-your-api-key"
)
response = client.chat.completions.create(
model="gemini-2.5-flash-lite",
messages=[
{"role": "system", "content": "You are an experienced software engineer."},
{"role": "user", "content": "Describe the architecture of a distributed cache."}
],
temperature=0.7,
max_tokens=1500
)
print(response.choices[0].message.content)