TL;DR (October 10, 2026): Gemini 3.8 Flash (gemini-3.8-flash, generally available since September 2, 2026) is free of charge on the Gemini API Free Tier for input, output and context caching, within your project's rate limits. On the Paid Tier it costs $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. Batch and Flex are half price, Priority is $1.35 and $6.75. Gemini 3.6 Flash has the same prices and the same January doubling. Free Tier rate limits are not published as a table; Google shows them per project in AI Studio, and daily quotas reset at midnight Pacific.
Google's Flash line moved fast in 2026: 3.6 Flash in July, 3.7 Flash in August, 3.8 Flash in September, and on October 8 Google deprecated 3.7 Flash and began routing its traffic to 3.8 Flash automatically. Do not confuse it with the original Gemini 3 Flash Preview, which is still on the price list at $0.50 / $3.00 but is now labelled a legacy model. 3.8 Flash is the current stable Flash model and the one most new projects should budget for, and its price comes with an end date, which matters for anyone budgeting into 2027.
Gemini 3.8 Flash pricing table
| Gemini 3.8 Flash, per 1M tokens | Free Tier | Paid, through Dec 31, 2026 | Paid, from Jan 1, 2027 |
|---|---|---|---|
| Standard input | Free of charge | $0.75 | $1.50 |
| Standard output (including thinking tokens) | Free of charge | $3.75 | $7.50 |
| Context caching (cached input) | Free of charge | $0.075 | $0.15 |
| Cache storage, per 1M tokens per hour | Not listed | $0.50 | $1.00 |
| Batch and Flex input / output | Not available | $0.375 / $1.875 | $0.75 / $3.75 |
| Priority input / output | Free of charge | $1.35 / $6.75 | $2.70 / $13.50 |
| Grounding with Google Search | Not available | 5,000 free requests a month (shared across Gemini 3 and newer), then $14 per 1,000 | |
| Content used to improve Google's products | Yes | No | |
From Google's Gemini Developer API pricing page (last updated October 9, 2026), read October 10, 2026. USD.
Every paid line doubles on January 1, 2027. There is no partial step. If you plan a budget for next year on today's rate, you will be off by exactly half.
How it compares with other Gemini Flash models
| Model | Free Tier | Paid input / output per 1M | Google's description |
|---|---|---|---|
| Gemini 3.8 Flash | Free | $0.75 / $3.75 (to Dec 31), then $1.50 / $7.50 | "Our most intelligent Flash model" |
| Gemini 3.6 Flash | Free | $0.75 / $3.75 (to Dec 31), then $1.50 / $7.50 | "Our previous generation Flash model" |
| Gemini 3.5 Flash-Lite | Free | $0.30 / $2.50 | Cost-efficient, high-volume agentic tasks |
| Gemini 3.1 Flash-Lite | Free | $0.25 / $1.50 (text, image, video) | Cost-efficient, high-volume tasks |
| Gemini 3 Flash Preview | Free | $0.50 / $3.00 (text, image, video) | "Our legacy Flash model" |
| Gemini 3.1 Pro Preview | Not available | $2.00 / $12.00 (prompts up to 200K) | Third-generation Pro model |
Two things stand out. First, 3.8 Flash costs the same as 3.6 Flash, so there is no price reason to stay on the older model. Second, the Flash-Lite models have no dated price change on the pricing page. From January, 3.5 Flash-Lite's $0.30 input is one fifth of 3.8 Flash's $1.50, which makes routing simple tasks to Flash-Lite much more attractive in 2027 than it is today. We compared Flash-Lite against non-Google models in the cheapest agentic LLM APIs.
Free Tier limits: where the numbers are
Google no longer prints a per-model Free Tier table in its rate limit docs. The page says limits "depend on a variety of factors (such as your usage tier) and can be viewed in Google AI Studio", and adds that "specified rate limits are not guaranteed and actual capacity may vary". What the docs do state:
- Three dimensions: requests per minute (RPM), input tokens per minute (TPM) and requests per day (RPD). Going over any one of them returns a rate-limit error, even if the other two have room.
- Per project, not per key. Extra API keys in the same project do not add capacity.
- RPD resets at midnight Pacific time.
- Preview and experimental models get tighter limits than stable ones. 3.8 Flash is a stable model.
- The Free Tier needs only an active project. No billing account is required.
If you use Gemini through a tool rather than directly, the tool may add its own cap. Gemini CLI, for example, gives an unpaid API key 250 requests a day on the Flash model, covered in our Gemini CLI free tier limits page.
Paid tier limits that catch people out
| Usage tier | Qualification | Monthly billing cap | Spend rate limit per 10 minutes |
|---|---|---|---|
| Free | Active project or free trial | N/A | N/A |
| Tier 1 | Active billing account linked | $250 | $10 |
| Tier 2 | $100 paid and 3 days since first payment | $2,000 | $50 |
| Tier 3 | $1,000 paid and 30 days since first payment | $20,000 to $100,000+ | $200 |
The spend rate limit is evaluated over a rolling 10-minute window and returns 429 RESOURCE_EXHAUSTED. At today's 3.8 Flash price, Tier 1's $10 window is about 13.3M input tokens in 10 minutes; from January it halves to about 6.7M. A batch backfill or a parallel agent run can hit that long before any token-per-minute limit. Other limits for Tier 1: Priority traffic gets 0.3x the standard rate limit, and the Batch API can hold 3,000,000 enqueued tokens for 3.8 Flash. The monthly caps are covered in Gemini API spend caps and tiers.
Worked cost examples
A support assistant: 1,000 requests a day, 10K input and 1.5K output tokens each. That is 10M input and 1.5M output a day.
- Through December 31, 2026: $7.50 + $5.63 = $13.13 a day, about $394 a month.
- From January 1, 2027: $15.00 + $11.25 = $26.25 a day, about $788 a month.
- Same work on Batch today: about $6.56 a day, if the replies can wait.
A large shared document with caching: 1,000 requests that each reuse a 100K-token cached document. Reading it uncached costs 100M × $0.75 = $75. Reading it from cache costs 100M × $0.075 = $7.50, plus storage of $0.05 per hour for 100K tokens, or $0.40 for an 8-hour working day. Caching saves about 90% here, as long as the requests actually hit the cache.
Remember that thinking tokens are billed as output. 3.8 Flash supports thinking levels low, medium and high; minimal is not supported and returns an error. On short-answer workloads, the thinking level is often the biggest lever on the output bill.
Model facts and migration notes
- Model code:
gemini-3.8-flash(stable). - Inputs: text, image, video, audio and PDF. Output: text.
- Token limits: 1,048,576 input, 65,536 output.
- Supported: caching, code execution, function calling, structured outputs, Search and Maps grounding, URL context, file search, computer use (preview), Batch, Flex and Priority. Not supported: Live API, audio or image generation.
- 3.7 Flash is deprecated (October 8, 2026) and its requests route to 3.8 Flash automatically. 3.5 Flash requests route to 3.6 Flash.
- Gemini 2.5 models have been limited since September 18, 2026 to users who used them before; Google recommends 3.5 Flash-Lite or 3.8 Flash for new projects.
- Sampling parameters
temperature,top_pandtop_kwere deprecated on July 21, 2026.
Free Tier or Paid Tier for production?
The Free Tier is fine for prototypes, but two points rule it out for most customer-facing work. Content sent on the Free Tier is used to improve Google's products, and the limits are not guaranteed. A common setup is a free project for experiments and a separate paid project for production, which our Gemini API free tier vs paid tier guide explains step by step.
Whichever you choose, record your own usage per customer and per feature, especially before January. When the price doubles, you will want to know which customers' margins it erases. UsageBox meters the usage events your application emits for that purpose; it does not import Google's billing data.
Sources
Read on October 10, 2026: the Gemini Developer API pricing page, the Gemini 3.8 Flash model page, the Gemini API rate limits page (last updated October 9, 2026), and the Gemini API changelog entries for July 21, September 2, September 18 and October 8, 2026.