TL;DR (updated September 13, 2026): Gemini API now has a clear Free Tier and Paid Tiers. A project on the Free Tier can use eligible models at the published free rate limits. A project upgraded to a Paid Tier is billed at the paid prices for its usage; it does not keep a hidden free allowance inside that paid project. Paid access now also depends on the billing plan attached to the billing account: many users are on Prepay, where you buy credits in advance, while eligible accounts can use Postpay. Rate limits are per project, billing tiers and account caps are determined at the billing-account level, and RPD resets at midnight Pacific.
Gemini API Free Tier vs Paid Tier in 2026
The simplest way to reason about Gemini API billing is to treat Free and Paid as different project states rather than as a single freemium meter.
| Free Tier project | Paid Tier project | |
|---|---|---|
| Model access | Eligible free-tier models only | Paid-tier and advanced model access according to current availability |
| Inference price | Free of charge where the model pricing table shows Free Tier | Published paid price from the first billable token/request |
| Rate limits | Free-tier limits | Higher limits based on billing tier |
| Data use | Free-service terms apply | Google states prompts/responses are not used to improve Google products under paid-service terms |
| Billing account | Not required | Required; Prepay or eligible Postpay |
Google's current pricing tables make the distinction explicit: individual models have separate Free Tier and Paid Tier columns. If a model is free on the Free Tier, the paid column still has its paid token/request price. If you want genuinely free experimentation alongside paid production, keep those workloads in separate projects rather than assuming the production project's first N requests will be free.
How Paid Tier setup works now
Moving a project to Paid Tier means linking it to an active Cloud Billing account and completing the billing setup in Google AI Studio. Google currently requires a minimum $5 prepayment when an account is assigned to Prepay. Some eligible accounts can use Postpay.
Paid tiers are based on payment history:
| Usage tier | Qualification | Billing-account tier cap |
|---|---|---|
| Tier 1 | Active billing account | $250/month |
| Tier 2 | $100 paid + 3 days since first successful payment | $2,000/month |
| Tier 3 | $1,000 paid + 30 days since first successful payment | $20,000–$100,000+/month |
Google says tiers, associated rate limits and billing-account caps are determined at the billing-account level. Projects linked to the same billing account inherit that account's tier. The account-level spend cap is also shared: when the cumulative Gemini API spend reaches the tier cap, service is paused for all linked projects until the next billing cycle unless an increase is approved.
Prepay changes the failure mode
For Prepay accounts, you buy credits first and Gemini usage deducts from that balance. New users default to Prepay as Google's rollout progresses. Credits expire after 12 months, and Google supports optional auto-reload.
The important operational detail is that a prepay balance is shared by projects under that billing account. When the balance hits zero, Gemini API service for all projects using that Prepay billing account stops. Google also documents roughly ten minutes of billing-pipeline latency for some long-running workloads, so agents and batch jobs can overshoot the nominal balance before the stop is observed.
Prepay has a separate monthly auto-charge limit for automatic reloads. That control limits automatic top-ups; it does not make manually purchased credits impossible.
Gemini spend caps are real, but not instantaneous
Gemini API now has two layers of spend-cap controls:
- Billing-account tier cap: the mandatory monthly ceiling associated with Tier 1, 2 or 3.
- Project spend cap: an experimental per-project control you can configure in AI Studio.
Do not design a synchronous authorization system around either number. Google says project spend-cap processing can lag by around ten minutes, and long-running batch or agent workloads may exceed the cap before billing data catches up. Treat these controls as billing guardrails, not request-by-request admission checks.
RPM, TPM and RPD are separate from billing caps
Gemini rate limits are commonly expressed as:
- RPM — requests per minute
- TPM — input tokens per minute
- RPD — requests per day
Exceeding any active dimension can trigger a rate-limit error even when you are comfortably below the other limits. Limits vary by model and tier, especially for preview/experimental models, so a static table in an article will go stale quickly. The live values for your account belong in Google AI Studio and the official rate-limit documentation.
Two rules are stable enough to build around: rate limits are applied per project, not per API key, and RPD resets at midnight Pacific. Creating more keys inside one project does not multiply its quota.
Can you keep free testing next to paid production?
Yes. Keep a Free Tier project unlinked from paid billing for eligible experimentation, and use a separate Paid Tier project for production. Google also documents that disabling billing on a paid project returns that project to the Free Tier, subject to current free-model availability and limits.
Separating projects has another benefit: it makes credentials and environments explicit. A production key cannot silently consume the testing project's quota, and a prototype cannot accidentally consume the production billing account's paid budget.
What to monitor in production
- Provider-side billing state. Know whether the billing account is Prepay or Postpay and whether auto-reload is enabled.
- Account and project spend caps. Set project caps where available, but budget for enforcement latency.
- Rate-limit headroom. Track RPM/TPM/RPD independently from dollar spend.
- Your own billable workload. Record the customer/workload quantity when your application performs the work, rather than relying only on a provider invoice weeks later.
- Reconcile. Compare your own period totals with the provider bill and investigate the gap: retries, provider-side tools, cache behavior, unmetered background jobs or missing application events.
Where UsageBox fits—and where it does not
UsageBox does not import Google AI Studio billing data, inspect your Gemini invoice, change Gemini tiers, or enforce Google's caps. Google remains the source of truth for the provider bill.
UsageBox is useful on the other side of the reconciliation: when your application can emit the usage event itself. You can preserve the customer/product/meter quantity independently, then compare that ledger with the provider invoice. That gives you evidence for questions such as "which customer workload produced this cost?" without pretending the application meter is the provider's billing system.
See the AI API billing reconciliation pattern →