Gemini 3.8 Flash Pricing & Free Tier (October 2026)

Gemini 3.8 Flash is free on the Gemini API Free Tier and costs $0.75 input and $3.75 output per 1M tokens through December 31, 2026, doubling to $1.50 and $7.50 on January 1, 2027. Batch, Flex and Priority prices, free tier rate limits, spend limits, and cost examples.

9 min read

Gemini 3.8 FlashGemini API pricingGemini free tierrate limitsAI cost

TL;DR (October 10, 2026): Gemini 3.8 Flash (gemini-3.8-flash, generally available since September 2, 2026) is free of charge on the Gemini API Free Tier for input, output and context caching, within your project's rate limits. On the Paid Tier it costs $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. Batch and Flex are half price, Priority is $1.35 and $6.75. Gemini 3.6 Flash has the same prices and the same January doubling. Free Tier rate limits are not published as a table; Google shows them per project in AI Studio, and daily quotas reset at midnight Pacific.

Google's Flash line moved fast in 2026: 3.6 Flash in July, 3.7 Flash in August, 3.8 Flash in September, and on October 8 Google deprecated 3.7 Flash and began routing its traffic to 3.8 Flash automatically. Do not confuse it with the original Gemini 3 Flash Preview, which is still on the price list at $0.50 / $3.00 but is now labelled a legacy model. 3.8 Flash is the current stable Flash model and the one most new projects should budget for, and its price comes with an end date, which matters for anyone budgeting into 2027.

Gemini 3.8 Flash pricing table

Gemini 3.8 Flash, per 1M tokensFree TierPaid, through Dec 31, 2026Paid, from Jan 1, 2027
Standard inputFree of charge$0.75$1.50
Standard output (including thinking tokens)Free of charge$3.75$7.50
Context caching (cached input)Free of charge$0.075$0.15
Cache storage, per 1M tokens per hourNot listed$0.50$1.00
Batch and Flex input / outputNot available$0.375 / $1.875$0.75 / $3.75
Priority input / outputFree of charge$1.35 / $6.75$2.70 / $13.50
Grounding with Google SearchNot available5,000 free requests a month (shared across Gemini 3 and newer), then $14 per 1,000
Content used to improve Google's productsYesNo

From Google's Gemini Developer API pricing page (last updated October 9, 2026), read October 10, 2026. USD.

Every paid line doubles on January 1, 2027. There is no partial step. If you plan a budget for next year on today's rate, you will be off by exactly half.

How it compares with other Gemini Flash models

ModelFree TierPaid input / output per 1MGoogle's description
Gemini 3.8 FlashFree$0.75 / $3.75 (to Dec 31), then $1.50 / $7.50"Our most intelligent Flash model"
Gemini 3.6 FlashFree$0.75 / $3.75 (to Dec 31), then $1.50 / $7.50"Our previous generation Flash model"
Gemini 3.5 Flash-LiteFree$0.30 / $2.50Cost-efficient, high-volume agentic tasks
Gemini 3.1 Flash-LiteFree$0.25 / $1.50 (text, image, video)Cost-efficient, high-volume tasks
Gemini 3 Flash PreviewFree$0.50 / $3.00 (text, image, video)"Our legacy Flash model"
Gemini 3.1 Pro PreviewNot available$2.00 / $12.00 (prompts up to 200K)Third-generation Pro model

Two things stand out. First, 3.8 Flash costs the same as 3.6 Flash, so there is no price reason to stay on the older model. Second, the Flash-Lite models have no dated price change on the pricing page. From January, 3.5 Flash-Lite's $0.30 input is one fifth of 3.8 Flash's $1.50, which makes routing simple tasks to Flash-Lite much more attractive in 2027 than it is today. We compared Flash-Lite against non-Google models in the cheapest agentic LLM APIs.

Free Tier limits: where the numbers are

Google no longer prints a per-model Free Tier table in its rate limit docs. The page says limits "depend on a variety of factors (such as your usage tier) and can be viewed in Google AI Studio", and adds that "specified rate limits are not guaranteed and actual capacity may vary". What the docs do state:

  • Three dimensions: requests per minute (RPM), input tokens per minute (TPM) and requests per day (RPD). Going over any one of them returns a rate-limit error, even if the other two have room.
  • Per project, not per key. Extra API keys in the same project do not add capacity.
  • RPD resets at midnight Pacific time.
  • Preview and experimental models get tighter limits than stable ones. 3.8 Flash is a stable model.
  • The Free Tier needs only an active project. No billing account is required.

If you use Gemini through a tool rather than directly, the tool may add its own cap. Gemini CLI, for example, gives an unpaid API key 250 requests a day on the Flash model, covered in our Gemini CLI free tier limits page.

Paid tier limits that catch people out

Usage tierQualificationMonthly billing capSpend rate limit per 10 minutes
FreeActive project or free trialN/AN/A
Tier 1Active billing account linked$250$10
Tier 2$100 paid and 3 days since first payment$2,000$50
Tier 3$1,000 paid and 30 days since first payment$20,000 to $100,000+$200

The spend rate limit is evaluated over a rolling 10-minute window and returns 429 RESOURCE_EXHAUSTED. At today's 3.8 Flash price, Tier 1's $10 window is about 13.3M input tokens in 10 minutes; from January it halves to about 6.7M. A batch backfill or a parallel agent run can hit that long before any token-per-minute limit. Other limits for Tier 1: Priority traffic gets 0.3x the standard rate limit, and the Batch API can hold 3,000,000 enqueued tokens for 3.8 Flash. The monthly caps are covered in Gemini API spend caps and tiers.

Worked cost examples

A support assistant: 1,000 requests a day, 10K input and 1.5K output tokens each. That is 10M input and 1.5M output a day.

  • Through December 31, 2026: $7.50 + $5.63 = $13.13 a day, about $394 a month.
  • From January 1, 2027: $15.00 + $11.25 = $26.25 a day, about $788 a month.
  • Same work on Batch today: about $6.56 a day, if the replies can wait.

A large shared document with caching: 1,000 requests that each reuse a 100K-token cached document. Reading it uncached costs 100M × $0.75 = $75. Reading it from cache costs 100M × $0.075 = $7.50, plus storage of $0.05 per hour for 100K tokens, or $0.40 for an 8-hour working day. Caching saves about 90% here, as long as the requests actually hit the cache.

Remember that thinking tokens are billed as output. 3.8 Flash supports thinking levels low, medium and high; minimal is not supported and returns an error. On short-answer workloads, the thinking level is often the biggest lever on the output bill.

Model facts and migration notes

  • Model code: gemini-3.8-flash (stable).
  • Inputs: text, image, video, audio and PDF. Output: text.
  • Token limits: 1,048,576 input, 65,536 output.
  • Supported: caching, code execution, function calling, structured outputs, Search and Maps grounding, URL context, file search, computer use (preview), Batch, Flex and Priority. Not supported: Live API, audio or image generation.
  • 3.7 Flash is deprecated (October 8, 2026) and its requests route to 3.8 Flash automatically. 3.5 Flash requests route to 3.6 Flash.
  • Gemini 2.5 models have been limited since September 18, 2026 to users who used them before; Google recommends 3.5 Flash-Lite or 3.8 Flash for new projects.
  • Sampling parameters temperature, top_p and top_k were deprecated on July 21, 2026.

Free Tier or Paid Tier for production?

The Free Tier is fine for prototypes, but two points rule it out for most customer-facing work. Content sent on the Free Tier is used to improve Google's products, and the limits are not guaranteed. A common setup is a free project for experiments and a separate paid project for production, which our Gemini API free tier vs paid tier guide explains step by step.

Whichever you choose, record your own usage per customer and per feature, especially before January. When the price doubles, you will want to know which customers' margins it erases. UsageBox meters the usage events your application emits for that purpose; it does not import Google's billing data.

Sources

Read on October 10, 2026: the Gemini Developer API pricing page, the Gemini 3.8 Flash model page, the Gemini API rate limits page (last updated October 9, 2026), and the Gemini API changelog entries for July 21, September 2, September 18 and October 8, 2026.

Key Topics

  • •Gemini 3.8 Flash
  • •Gemini API pricing
  • •Gemini free tier
  • •rate limits
  • •AI cost

Related Articles

Explore more articles on similar topics to deepen your understanding of usage-based billing.

Free AI APIs 2026: No Credit Card + Real Rate Limits

Free AI APIs that still work without a credit card in 2026, including OpenRouter and Gemini, with the RPM, RPD and token...

10 min readRead more

Does a 429 Count Against Your Rate Limit? Yes - and It Can Cost You Tokens (2026)

A rejected 429 request still counts. On the major LLM APIs a call that comes back 429 Too Many Requests generally counts...

10 min readRead more

Gemini CLI Free Tier Limits 2026: 1,000 Requests a Day

Gemini CLI limits by sign-in method as of October 2026: 1,000 requests a day with a Google account, 250 a day (Flash onl...

8 min readRead more

Explore More Articles

Discover our complete collection of usage-based billing guides and implementation patterns.

View all articles