Skip to content
DigitalTatva
Pricing and Plans

Gemini API Pricing: Free Tier Limits and Paid Token Costs

Gemini API pricing for September 2026: free tier limits, paid rates for 3.8 Flash, Flash-Lite and 3.1 Pro, the 2027 Flash price rise and ways to save.

Gemini API Pricing: Free Tier Limits and Paid Token Costs
On this page
  1. Key takeaways
  2. Gemini API pricing table: paid Standard tier
  3. The January 2027 Gemini Flash price increase
  4. Gemini API free tier: what you get and the limits
  5. Paid usage tiers and billing caps
  6. Worked examples: what the Gemini API really costs
  7. Batch and Flex: half-price Gemini API
  8. Gemini API vs OpenAI and Claude on price
  9. Gemini API vs Google AI Pro subscription
  10. Which Gemini model should you use?
  11. Frequently asked questions
  12. Our verdict on Gemini API pricing

Gemini API pricing has two layers: a free tier that lets you use Gemini Flash and Flash-Lite models at no cost within rate limits, and a paid tier billed per million tokens. As of September 2026, Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, Gemini 3.1 Flash-Lite costs $0.25 and $1.50, and the preview Gemini 3.1 Pro costs $2 and $12 for prompts up to 200K tokens.

There is one date every Gemini developer needs in their calendar: on January 1, 2027, the 3.6, 3.7 and 3.8 Flash models double in price to $1.50 input and $7.50 output. This guide covers the free tier, every paid rate, that upcoming increase, caching storage fees, grounding costs and how to keep your bill low.

Key takeaways

  • The free tier covers Flash and Flash-Lite models at no charge, but Gemini 3.1 Pro is paid only.
  • Gemini 3.8 Flash is $0.75 / $3.75 per 1M tokens until December 31, 2026, then $1.50 / $7.50 from January 1, 2027.
  • Gemini 3.1 Pro (preview) is $2 / $12 up to 200K prompt tokens and $4 / $18 above that.
  • Batch and Flex cut every rate by 50%, and context caching reads cost a tenth of normal input.
  • Grounding with Google Search is free for 5,000 requests a month, then $14 per 1,000.

Gemini API pricing table: paid Standard tier

All prices below are USD per 1 million tokens, from Google’s official Gemini API pricing page. If you need a refresher on what a token is, read our plain English guide to tokens.

Model Input Output Context caching Cache storage (per 1M tokens per hour) Status
Gemini 3.1 Pro (up to 200K prompt) $2.00 $12.00 $0.20 $4.50 Preview
Gemini 3.1 Pro (over 200K prompt) $4.00 $18.00 $0.40 $4.50 Preview
Gemini 3.8 Flash $0.75 $3.75 $0.075 $0.50 Stable
Gemini 3.7 Flash $0.75 $3.75 $0.075 $0.50 Stable
Gemini 3.6 Flash $0.75 $3.75 $0.075 $0.50 Stable
Gemini 3.5 Flash $1.50 $9.00 $0.15 $1.00 Stable
Gemini 3.5 Flash-Lite $0.30 $2.50 Reported as not available n/a Stable
Gemini 3.1 Flash-Lite $0.25 ($0.50 audio) $1.50 Check pricing page Check pricing page Stable

A few details matter here. Gemini 3.5 Flash-Lite charges the same $0.30 for text, image, video and audio input, while 3.1 Flash-Lite charges $0.25 for text, image and video but $0.50 for audio. Gemini 3.5 Flash also has a Priority tier at $2.70 input and $16.20 output for faster, more reliable processing.

Bar chart of Gemini API output prices per million tokens: 3.1 Pro $18 over 200K and $12 up to 200K, 3.5 Flash $9, 3.8 Flash $3.75, 3.5 Flash-Lite $2.50, 3.1 Flash-Lite $1.50
Gemini API output price by model

The January 2027 Gemini Flash price increase

The current Flash prices come with an end date. Google’s pricing page shows that Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash keep their $0.75 / $3.75 rates until December 31, 2026. From January 1, 2027, all three move to:

  • Input: $1.50 per 1M tokens (up from $0.75)
  • Output: $7.50 per 1M tokens (up from $3.75)
  • Context caching: $0.15 per 1M tokens (up from $0.075)
  • Cache storage: $1.00 per 1M tokens per hour (up from $0.50)

In other words, every Flash rate doubles. If you are building a budget or pricing your own product on top of Gemini Flash, use the 2027 numbers now so you are not caught out in three months.

Watch out: The 2027 Flash price ($1.50 / $7.50) matches the current Gemini 3.5 Flash price exactly. Moving from 3.8 Flash to an older Flash model will not save money after the increase. Flash-Lite models are the cheaper escape route if quality holds up for your task.

Gemini API free tier: what you get and the limits

Google’s free tier is one of the best ways to start building with AI at no cost. According to the pricing page, the Flash and Flash-Lite models are “Free of charge” on the free tier. Gemini 3.1 Pro is not available on the free tier at all.

What Google no longer does is publish exact free tier rate limits in its docs. The rate limits documentation now tells you to check your limits inside Google AI Studio. A developer forum post from early September 2026 reported free limits of around 20 requests per day for Gemini 3.8 Flash and around 500 per day for Flash-Lite. That is a user report, not an official number, so check your own AI Studio dashboard.

If those figures are right, the free tier is fine for learning, prototypes and personal scripts, but not for a public app. Flash-Lite’s more generous allowance makes it the better choice for free experiments. For more zero-cost options, see our roundup of the best free AI tools.

Tip: Before you send customer or confidential data through the free tier, read Google’s current terms on how free tier data is handled. Paid tier terms can differ, and that alone can be a reason to link billing early.

Once you link a billing account, Google places your project in a usage tier. Higher tiers raise your rate limits and your billing cap. Qualification is based on how much you have paid and for how long.

Tier How to qualify Billing cap
Tier 1 Link a billing account $250
Tier 2 $100 paid and 3 days since first payment $2,000
Tier 3 $1,000 paid and 30 days since first payment $20,000 to $100,000+

The Tier 1 cap of $250 is a useful safety net for small projects. If you expect heavy traffic at launch, plan ahead: moving to Tier 2 takes both spend and at least three days.

Worked examples: what the Gemini API really costs

Here is the arithmetic on four typical workloads. Our guide to estimating AI API costs walks through the method in more detail.

Example 1: a chatbot with 1,000 conversations a day

Each conversation uses about 2,000 input tokens and 500 output tokens, so 2M input and 0.5M output tokens per day.

Model Calculation Per day Per 30 days
3.1 Pro (under 200K) (2 x $2) + (0.5 x $12) $10.00 $300
3.8 Flash (2026) (2 x $0.75) + (0.5 x $3.75) $3.38 about $101
3.8 Flash (from 2027) (2 x $1.50) + (0.5 x $7.50) $6.75 about $203
3.1 Flash-Lite (2 x $0.25) + (0.5 x $1.50) $1.25 about $38

Example 2: the Gemini 3.1 Pro 200K threshold

A single 300,000-token prompt with a 5,000-token answer on 3.1 Pro crosses the 200K line, so it is billed at the higher rates: (0.3 x $4) + (0.005 x $18) = $1.20 + $0.09 = $1.29. If you could split it into two prompts under 200K, the same tokens at the lower rate would cost (0.3 x $2) + (0.005 x $12) = $0.60 + $0.06 = $0.66. Our context window explainer covers how to chunk long documents.

Example 3: context caching a large document

You load a 100,000-token product manual and ask 200 questions about it over 8 hours on 3.8 Flash.

  • Without caching: 200 x 0.1M x $0.75 = $15.00 in input.
  • With caching: reads cost 200 x 0.1M x $0.075 = $1.50, plus storage of 0.1M x $0.50 x 8 hours = $0.40, for about $1.90.

Unlike some rivals, Gemini charges an hourly storage fee for cached content, so caching pays off when a cache is read many times, and it can cost more than it saves if you store a big cache and barely use it. Our prompt caching guide compares how each provider handles this.

Example 4: a grounded research assistant

Your app makes 20,000 requests a month with Grounding with Google Search. The first 5,000 are free, and the other 15,000 cost 15 x $14 = $210, on top of token costs. Grounding is powerful but it can outweigh token spend for search-heavy apps.

Batch and Flex: half-price Gemini API

Google’s Batch and Flex options are both 50% off Standard rates. Gemini 3.8 Flash in batch mode costs $0.375 input and $1.875 output, and 3.5 Flash-Lite costs $0.15 and $1.25. Batch suits jobs that do not need an instant answer: bulk classification, nightly summaries, dataset labeling and content rewrites.

Example: summarizing 5,000 reports at 4,000 tokens in and 400 out is 20M input and 2M output tokens. On 3.8 Flash Standard: (20 x $0.75) + (2 x $3.75) = $15 + $7.50 = $22.50. In batch mode: $11.25. Read our Batch API guide for a cross-vendor comparison.

Gemini API vs OpenAI and Claude on price

Model Input / 1M Output / 1M Free tier
Gemini 3.1 Flash-Lite $0.25 $1.50 Yes
Gemini 3.8 Flash (2026 price) $0.75 $3.75 Yes
OpenAI GPT-6 Luna $0.10 $0.50 Not confirmed
Claude Haiku 4.5 $1.00 $5.00 No
Gemini 3.1 Pro (under 200K) $2.00 $12.00 No
OpenAI GPT-6.1 Sol $2.00 $10.00 Not confirmed
Claude Sonnet 5.5 $2.00 $10.00 No

Gemini’s biggest advantage is the free tier: no other major lab clearly offers free production-grade models for testing. On paid list price, OpenAI’s GPT-6 Luna undercuts Gemini’s budget models, and Gemini 3.1 Pro’s output rate is slightly above GPT-6.1 Sol and Claude Sonnet 5.5. See our OpenAI API pricing guide and Claude API pricing guide for full rate cards, or how to pick the cheapest model for each task.

Gemini API vs Google AI Pro subscription

The Gemini API and Google’s consumer subscriptions are separate products. Google AI Pro costs $19.99 a month in the US and Rs 1,950 a month in India, and gives you the Gemini app, Gemini in Gmail and Docs, 5 TB of storage and more. It does not include API usage. Developers pay for the API through Google AI Studio or Google Cloud billing. For the consumer plans, see our Gemini pricing guide, and for everyday office use, our guide to using Gemini in Gmail, Docs and Sheets.

Which Gemini model should you use?

  • Students and hobbyists: the free tier with Flash-Lite, then Flash for harder prompts.
  • Startups building chat or content features: Gemini 3.8 Flash, but budget at the 2027 price.
  • High-volume classification or extraction: 3.1 Flash-Lite or 3.5 Flash-Lite, ideally through Batch.
  • Complex reasoning and coding: Gemini 3.1 Pro, keeping prompts under 200K tokens where possible. Remember it is still a preview model, so test for stability.
  • Research tools that need fresh facts: any Flash model with grounding, watching the $14 per 1,000 fee after 5,000 free requests.

Pros

  • Genuine free tier for Flash and Flash-Lite models
  • Low Flash-Lite pricing for bulk work
  • Batch and Flex at 50% off
  • 5,000 free grounded searches per month

Cons

  • Flash prices double on January 1, 2027
  • No stable Pro model, only a preview
  • Hourly cache storage fees add up
  • Free tier limits are no longer published in the docs

Frequently asked questions

Is the Gemini API free?

Partly. Google’s free tier lets you use Gemini Flash and Flash-Lite models free of charge within rate limits shown in Google AI Studio. Gemini 3.1 Pro is not on the free tier. The free limits are fine for learning and prototypes, but a public app will need a billing account and the paid rates, starting with Tier 1 and its $250 billing cap.

What are the Gemini API free tier limits?

Google no longer lists per-model free limits in its documentation and asks you to check them in AI Studio. A September 2026 developer forum post reported around 20 requests per day for Gemini 3.8 Flash and around 500 per day for Flash-Lite. Treat those as unofficial and confirm the limits shown for your own project.

How much does Gemini 3.1 Pro API cost?

Gemini 3.1 Pro, which is still in preview, costs $2 per million input tokens and $12 per million output tokens for prompts up to 200K tokens. Longer prompts are billed at $4 input and $18 output. Context caching is $0.20 or $0.40 per million depending on prompt size, plus $4.50 per million tokens per hour for storage.

Is the Gemini API cheaper than OpenAI?

It depends on the tier. Gemini offers a free tier OpenAI does not clearly match, and 3.1 Flash-Lite is inexpensive at $0.25 / $1.50. But OpenAI’s GPT-6 Luna is cheaper still at $0.10 / $0.50. At the mid tier, Gemini 3.1 Pro and GPT-6.1 Sol share a $2 input price, while Gemini charges $12 output versus $10.

How much does the Gemini API cost in India?

Google lists Gemini API rates in US dollars, and we found no separate rupee rate card for API tokens. Indian developers pay the USD rates through their Google billing account, so the rupee cost depends on exchange rates and applicable taxes. The free tier works the same in India, which makes it a good starting point for students and early-stage builders.

Our verdict on Gemini API pricing

The Gemini API is the easiest place to start for free, and its Flash-Lite models are among the cheapest ways to run bulk AI work. Gemini 3.8 Flash is excellent value right now, but its price doubles on January 1, 2027, so plan for $1.50 / $7.50 in any long-term budget. For heavy reasoning, 3.1 Pro is competitive if you stay under the 200K prompt threshold.

Use Batch for anything that can wait, cache only what you will reuse often, and keep an eye on grounding fees. Google changes rates and limits regularly, so confirm the current numbers on the official Gemini API pricing page before you commit.

Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.

Written by

Ketan Parmar

Ketan Parmar has spent more than 15 years in digital marketing, helping brands grow through SEO, Google Ads, Meta Ads, content strategy and social media. Today he focuses on AI search visibility: how businesses get found and recommended in ChatGPT, Gemini, Perplexity and Google's AI answers.

Get the weekly AI tools brief

New tools, price changes and money-saving deals. One email a week, no spam.