Gemini API pricing runs $0.75 to $18.00 per million tokens across the current lineup, collected from Google's own documentation on September 11, 2026.
The cheapest advertised rate will not last. Every 3.x Flash price doubles on January 1, 2027, and Grounding with Google Search bills per search query on Gemini 3, not per prompt.
Check the billing section before you enable it.
Model | Input, per 1M tokens | Output, per 1M tokens | Notes |
|---|---|---|---|
Gemini 3.8 Flash ( | $0.75 → $1.50 on 2027-01-01 | $3.75 → $7.50 on 2027-01-01 | Free tier available. Batch: 50% off. |
Gemini 3.7 Flash ( | $0.75 → $1.50 | $3.75 → $7.50 | Same schedule as 3.8 Flash. |
Gemini 3.6 Flash ( | $0.75 → $1.50 | $3.75 → $7.50 | Same schedule as 3.8 Flash. |
Gemini 3.1 Pro Preview ( | $2.00 (≤200K), $4.00 (>200K) | $12.00 (≤200K), $18.00 (>200K) | No introductory-rate expiry stated. |
Gemini 2.5 Pro ( | $1.25 (≤200K), $2.50 (>200K) | $10.00 (≤200K), $15.00 (>200K) | Grounding: 1,500 free requests/day, then $35/1,000. |
Gemini 2.5 Flash ( | $0.30 (text), $1.00 (audio) | $2.50 | No 200K-token price step. |
This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed. Every figure below comes from Google's own pricing and billing documentation, retrieved September 11, 2026, plus the arithmetic shown where a figure is derived rather than quoted directly. No figure here was produced by testing. The full source list with retrieval dates lives on the methodology page.
Seven decisions move the real bill away from the rate-card number above. They are ranked here by size of impact: the January 2027 Flash increase, how grounding is actually counted, what enabling billing does to the free allowance, the current Flash and Pro roster, the Flex and Priority tiers next to Batch, and the costs that live outside token pricing entirely.
This is one spoke of a wider API pricing comparison covering every current model tracked here, not only Gemini.

The Flash rate on today's pricing page expires December 31, 2026
Gemini 3.8 Flash, Gemini 3.7 Flash, and Gemini 3.6 Flash bill at $0.75 per 1 million input tokens and $3.75 per 1 million output tokens through December 31, 2026. That changes on January 1, 2027.
Google's pricing documentation states the same rate rises to $1.50 input and $7.50 output that day, a flat doubling on both sides of the ledger. Context caching moves with it.
The cached-read rate rises from $0.075 to $0.15 per 1 million tokens, and cache storage rises from $0.50 to $1.00 per 1 million tokens per hour, both on the same date.
A workload budgeted today at the current Flash rate is underpriced for any month after December 2026. Not by some vague margin. By exactly 100 percent, on both input and output.
Nobody reading a rate card in September has a reason to expect a full doubling four months out unless the date sits next to the number, which is why it appears here rather than in a footnote.

Grounding with Google Search bills by the search, not by the prompt
On Gemini 3 models, Google's grounding documentation states the billing unit directly: a project "is billed for each search query that the model decides to execute." One API call is not one billable event. If the model runs five searches to answer a single prompt, that call produces five billable search queries at $14 per 1,000. The rule changes for older models.
The same documentation states it plainly: "This billing model only applies to Gemini 3 models; when you use search grounding with Gemini 2.5 or older models, your project is billed per prompt." Two different billing systems sit inside one product line, split by model generation.
The free allowance for Gemini 3.x grounding is 5,000 search requests per month. That pool is shared across every Gemini 3.x model a project calls, not issued separately per model.
Gemini 2.5 Pro grounds on a different meter entirely: 1,500 requests per day free, then $35 per 1,000 grounded prompts, a per-prompt rate more than double the per-query Gemini 3 rate.
An agentic workflow that fans out to three or four searches per turn on Gemini 3 spends its 5,000-query pool three or four times faster than a single-search estimate suggests. The paid overage compounds the same way once the pool is gone.

Enabling billing does not carry the free allowance into the paid tier
Google's billing documentation draws one line clearly. On the free tier, prompts and responses are used to improve Google's products.
On paid usage, the documentation states plainly that prompts and responses "are not used to improve Google products." That distinction is confirmed and dated.
What the documentation leaves unresolved is whether a project's free-tier request allowance persists once billing is linked. Developer reports describe the opposite of what several third-party guides assume: the free quota does not sit underneath the paid tier as a standing buffer, the way it does on some other Google Cloud services.
A project with billing turned on should be budgeted as fully metered from its first paid token, not as free-quota-plus-overage.
Spend caps exist because that gap has produced billing surprises. Google's billing account tiers carry system-level monthly caps: $250 on Tier 1, $2,000 on Tier 2, and $20,000 to $100,000 or more on Tier 3.
A project owner can also set a custom monthly spend cap inside AI Studio's Spend tab. Both cap types share one mechanical limit. Enforcement runs on roughly a 10-minute delay, and Google's own documentation states the developer is responsible for usage that lands inside that window. Once a cap is hit, the project pauses until the first of the next month, not until the developer manually resets it.
The current Flash generation is 3.6, 3.7, and 3.8, not 2.0 or 1.5
Three Flash models are current as of September 11, 2026: Gemini 3.6 Flash, Gemini 3.7 Flash, and Gemini 3.8 Flash. All three price identically at $0.75 input and $3.75 output through the end of 2026. Gemini 2.0 Flash and Gemini 1.5 Pro are retired lines, not budget alternatives to the 3.x family, and neither carries a current rate on Google's own pricing page.
Two Pro-tier models are current: Gemini 3.1 Pro Preview, still labeled Preview on Google's documentation, and Gemini 2.5 Pro, priced separately and lower at every tier.
Gemini 2.5 Flash sits alongside the 3.x Flash family at a different structure entirely. Text, image, and video input price at $0.30 per 1 million tokens; audio input prices at $1.00; output prices at $2.50, with no 200,000-token step.
A reader comparing 2.5 Flash against 3.6, 3.7, or 3.8 Flash on price alone is comparing a cheaper input rate against a cheaper output rate, not a strictly cheaper or strictly pricier model.
Flex halves the price without the 24-hour wait Batch requires
Four processing tiers sit under the same per-model token rate, and they trade speed for discount in different directions.
Standard is the baseline rate shown in the table above. Batch runs asynchronously, with a turnaround of up to 24 hours, at 50 percent of the Standard rate.
Flex runs synchronously at that same 50 percent discount, targeting a 1 to 15 minute response window rather than Batch's 24 hours.
That window makes Flex usable for a workflow where one request's output feeds the next.
Flex requests are also the first to go. Google's Flex inference documentation states Flex traffic is treated with lower priority and can be preempted or evicted when Standard demand spikes, returning a 503 or 429 error, with no automatic upgrade to Standard when that happens.
Priority sits at the opposite end: 75 to 100 percent above Standard, for latency measured in seconds, and reliability the documentation describes as non-sheddable.
Flex is the only half-price option that keeps a synchronous, sub-15-minute response. It is also the only one of the four tiers that can silently drop a request under load.
A reader sizing a production workload around the 50 percent discount alone, without checking which of the two half-price tiers they landed on, finds that out the hard way.

Spend caps, Vertex AI, and the costs that live outside token pricing
Two costs sit next to the per-token rate rather than inside it. The first is regional routing. Vertex AI's own pricing documentation prices non-global Gemini endpoints at a 10 percent premium over the global endpoint rate, as of July 1, 2026.
That surcharge exists on Vertex AI specifically and has no equivalent inside the Gemini Developer API, which routes globally by default.
The two products otherwise price the same current models at the same base per-token rate. Vertex AI adds the regional premium plus its own enterprise governance layer, and neither product publishes a bundled discount for choosing one over the other.
The second cost is audio. Gemini 2.5 Flash prices audio input at $1.00 per 1 million tokens against $0.30 for text, on the identical model.
That is a more than three-fold difference a reader checking only the headline text rate never sees.
A team pricing a voice application off the text row in any rate table is pricing the wrong column.
What this article did not measure
This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed. It did not independently verify Google's billing-tier spend-cap dollar amounts against every regional billing configuration, since Google's own documentation states the figures as standard tiers without listing regional exceptions.
It did not confirm whether the 5,000-query monthly grounding pool is scoped per Google Cloud project or aggregated per billing account across multiple projects.
Google's own developer forum shows this question asked and unanswered as of the collection date below, so it stays open here rather than guessed.
Which Gemini tier fits which reader
Developers
Best value: Gemini 3.6, 3.7, or 3.8 Flash on the free tier, for prototyping. Free-tier prompts are used to improve Google's products, a trade a solo developer testing a prototype is more likely to accept than a team handling user data.
Startups
Situational: Flex processing on a current Flash model covers a cost-sensitive production workload that can tolerate a 1 to 15 minute response and an occasional dropped request, at half the Standard rate.
A workload that cannot tolerate a dropped request belongs on Standard or Priority instead.
Enterprise
Situational: Vertex AI's 10 percent regional premium buys data residency and governance controls the Gemini Developer API does not offer. That trade is worth making only when a regional requirement already exists, not as a default.
High volume
Top pick: Batch processing, at 50 percent off Standard, for any workload where a 24-hour turnaround is acceptable. This removes the Flex tier's shedding risk entirely.
Low latency
Situational: Priority processing, at 75 to 100 percent above Standard, is the only tier Google's documentation describes as non-sheddable. That matters more than the premium for a workload that cannot silently drop a request.
On-device
Not applicable. The Gemini API prices cloud-hosted inference; this article priced no on-device or locally-run Gemini variant.
Self-hosting
Not applicable. Google does not publish open-weight Gemini models, so no self-hosting cost comparison exists for this product line.
Non-English
Insufficient data. None of the pricing or billing documentation reviewed for this article publishes a per-language rate or a non-English usage adjustment.
Is the Gemini API free?
Gemini offers a genuine free tier across its Flash models, with one standing condition: free-tier prompts and responses are used to improve Google's products. Gemini 3.1 Pro Preview and Gemini 2.5 Pro carry no equivalent free allowance in the documentation reviewed here.
Enabling a billing account moves a project onto metered, non-training-use pricing. Budget it as fully paid from the first token, not as free-quota-plus-overage.
How does Gemini pricing compare to GPT and Claude?
On September 11, 2026, OpenAI's own pricing page lists GPT-5.6 Sol at $5.00 input and $30.00 output per 1 million tokens, and GPT-5.6 Luna at $0.20 input and $1.20 output. Against that field, Gemini 3.8 Flash's $0.75 and $3.75 undercuts every budget-tier model listed except GPT-5.6 Luna's output rate.
Anthropic's own pricing documentation lists Claude Sonnet 5 at $2.00 input and $10.00 output, and Claude Opus 5 at $5.00 input and $25.00 output, with a full current-generation breakdown on the Claude API pricing page.
Gemini 3.1 Pro Preview's $2.00 to $4.00 input range sits below both flagship models' input rates, while its $12.00 to $18.00 output range sits between Claude Sonnet 5 and Claude Opus 5.

What is the cheapest Gemini model?
Gemini 3.6 Flash, Gemini 3.7 Flash, and Gemini 3.8 Flash tie as the cheapest current models, at $0.75 input and $3.75 output per 1 million tokens, through December 31, 2026.
Gemini 2.5 Flash prices lower on text input, at $0.30, but higher on audio input, at $1.00. Which model is actually cheaper depends on the input type a workload sends, not on one headline number.
Is the Gemini Developer API cheaper than Vertex AI?
The two products price identical current models at identical base per-token rates. Vertex AI adds a 10 percent premium specifically on non-global regional endpoints, a cost the Gemini Developer API does not carry since it routes globally by default.
A workload with no regional requirement pays the same either way. A workload that requires a specific region pays 10 percent more on Vertex AI for that requirement.
Does enabling billing delete my free tier allowance?
Google's own billing documentation does not describe a standing free-quota buffer inside the paid tier, the way some other Google Cloud products carry one.
Developer reports describe billing turning on a fully metered project, not a paid tier stacked on top of the prior free allowance.
Budget a project as though the free tier ends the moment a billing account links, not as a cushion against the first bill.
Will my Gemini bill go up on January 1, 2027 without me changing anything?
Yes, for any workload on Gemini 3.6, 3.7, or 3.8 Flash. Google's own pricing documentation states the input and output rates on all three double that day, from $0.75 and $3.75 to $1.50 and $7.50, with cached-read and cache-storage rates doubling alongside them.
No code change or usage change triggers it. The date alone does.
