Gemini CLI Free Tier Limits 2026: 1,000 Requests a Day

Gemini CLI limits by sign-in method as of October 2026: 1,000 requests a day with a Google account, 250 a day (Flash only) with an unpaid API key, 1,500 to 2,000 on paid plans. What counts as a request, which plans are not supported, and what pay-as-you-go costs.

8 min read

Gemini CLIGemini Code Assistfree tierrate limitsAI coding tools

TL;DR (October 10, 2026): Gemini CLI's free tier depends on how you sign in. Log in with Google (Gemini Code Assist for individuals) and you get 1,000 model requests per user per day across the Gemini model family. Use an unpaid Gemini API key and you get 250 requests per day, Flash model only. Vertex AI Express mode gives you 90 days before billing is required. Paid seats raise the daily cap to 1,500 (Google AI Pro, Code Assist Standard) or 2,000 (Google AI Ultra, Code Assist Enterprise, Workspace AI Ultra). One prompt can use several requests, and the cap covers all models combined.

"1,000 free requests a day" is the number everyone quotes for Gemini CLI, and it is correct, but only for one of the four ways to sign in. The other routes have different caps, different model access and different billing. The details are spread across the Gemini CLI quota page, Google Cloud's Gemini Code Assist quota page and the Gemini API rate limit docs. This page brings them together as of October 2026.

Gemini CLI limits by sign-in method

Sign-in methodTier or subscriptionMax requests per user per dayNotes
Google accountGemini Code Assist for individuals (free)1,000Requests go across the Gemini model family, as the CLI decides
Google AI Pro1,500Fixed-price personal subscription
Google AI Ultra2,000Fixed-price personal subscription
Gemini API keyFree tier (unpaid)250Flash model only
Pay-as-you-goVariesBilled per token by your Gemini API tier
Vertex AIExpress mode (free)Varies90 days before you need to enable billing
Pay-as-you-goVariesDynamic shared quota or provisioned throughput
Google WorkspaceCode Assist Standard1,500Licensed seat
Code Assist Enterprise2,000Licensed seat
Workspace AI Ultra2,000Workspace add-on

From the Gemini CLI "Quotas and pricing" page (last updated June 18, 2026) and Google Cloud's Gemini Code Assist quotas page (last updated October 9, 2026), both read October 10, 2026.

On top of the daily cap, the CLI docs say requests "are limited per user per minute and are subject to the availability of the service in times of high demand". So even with requests left for the day, a burst of activity or a busy period can slow you down.

What counts as one request

This is where most people misjudge the free tier. A request is one model call, not one prompt. Google's Code Assist quota page says it directly: "When in agent mode or when using the Gemini CLI, one prompt might result in multiple model requests." An agent that reads three files, runs a test and edits two files may make many model calls to answer a single instruction.

Three more rules from the same page:

  • All models share one cap. The daily limits "are aggregated across all interactions with any model version or family (for example, Pro, Flash)". There is no separate Pro and Flash allowance.
  • Agent mode and CLI share one cap. Requests from Gemini Code Assist agent mode in your IDE and from Gemini CLI are combined. Heavy IDE use in the morning leaves less for the terminal in the afternoon.
  • The CLI picks the model. With Google sign-in, requests go "across the Gemini model family as determined by Gemini CLI". You can choose a model, but the quota is the same pool.

In practice, 1,000 requests is generous for asking questions and making small edits, and much tighter for long autonomous runs. If you want a per-task view of how fast agent loops consume budget, our guide to measuring what one coding-agent task costs walks through it.

Which plans are not supported

Several Google subscriptions do not raise your Gemini CLI quota, and the docs name them:

  • Google AI Plus is not a supported tier for personal accounts. Only AI Pro and AI Ultra are.
  • Workspace AI Standard, Workspace AI Plus and AI Expanded are not supported for Workspace accounts.
  • Gemini for Workspace plans apply to Google's web products such as the Gemini app, "not to the API usage which powers the Gemini CLI". Google says support is "under active consideration".

To check whether you are signed in with a personal or a Workspace account, the docs suggest opening the Google One plans page: a Workspace account shows the message "You're currently signed in to your Google Workspace Account".

The API key route: 250 a day, Flash only

Signing in with an unpaid Gemini API key gives 250 requests a day and only the Flash model. That is a quarter of the Google sign-in allowance, so for interactive use the Google account route is better. The API key route matters for two other reasons:

  1. Once paid, it has no daily request ceiling. The CLI docs call pay-as-you-go through an API key or Vertex AI "the recommended path for uninterrupted access", for example when you "exhaust your Gemini Pro quota even after upgrading".
  2. It follows Gemini API rules. Gemini API rate limits are applied per project, not per API key, and requests-per-day quotas reset at midnight Pacific time. Specific free-tier numbers are shown in Google AI Studio rather than in the docs.

Free Gemini API usage also comes with different data terms: Google's pricing page says free-tier content is used to improve Google's products, and paid-tier content is not. We explain the Free and Paid split in detail in Gemini API billing: free tier vs paid tier.

What pay-as-you-go costs

Once you switch to a paid API key, you pay per token at the price of whichever model the CLI calls. As of October 2026, Google's pricing page lists Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026, rising to $1.50 and $7.50 from January 1, 2027.

Example: 1,000 requests a day, 50K input and 2K output tokens eachThrough Dec 31, 2026From Jan 1, 2027
Input (50M tokens)$37.50$75.00
Output (2M tokens)$7.50$15.00
Per day$45.00$90.00

The point of the example is the input column. CLI sessions resend a lot of context, so input tokens dominate. Thinking tokens are billed as output on Gemini models. The Gemini CLI docs also note that per-token billing "can be more expensive for many small calls with few tokens", which is why the fixed-price subscriptions exist.

Paid Gemini API projects can also hit spend-based rate limits over a rolling 10-minute window (Google says whether they apply depends on billing history and account standing): $10 on Tier 1, $50 on Tier 2 and $200 on Tier 3, returning 429 RESOURCE_EXHAUSTED when hit. A long agent run on a new Tier 1 account can trip the $10 window before it trips any token limit. Our Gemini API spend caps and tiers article covers the monthly caps that sit above this.

Checking your usage

  • /stats model shows current token usage for the session and the limits for your current quota.
  • A summary of model usage prints when you exit a session.
  • For API key usage, the AI Studio rate-limit dashboard shows your project's active limits.

Gemini CLI is open source and moves quickly: the latest stable release on GitHub at the time of writing is v0.63.0, published October 6, 2026, with preview and nightly builds after it. Quota behavior is set on Google's side, not in the client version.

Which option to pick

  • Learning or occasional use: sign in with a personal Google account and use the free 1,000 requests.
  • Daily professional use, predictable cost: Google AI Pro (1,500) or Ultra (2,000), or a Code Assist Standard or Enterprise seat through your company.
  • CI, scripts, or long unattended runs: a paid Gemini API key or Vertex AI, with a spend cap set, because there is no daily stop.

If you run the CLI for a team on a shared paid key, the Google bill shows the project total only. Recording which person or pipeline made each call is up to you, and that is the part UsageBox handles: it meters the usage events your own scripts send, and does not read Google's billing data.

Sources

Read on October 10, 2026: the Gemini CLI quotas and pricing page (also in the google-gemini/gemini-cli repository on GitHub), Google Cloud's Gemini Code Assist quotas and limits page, the Gemini API rate limits page, the Gemini Developer API pricing page, and the Gemini CLI releases list on GitHub.

Key Topics

  • •Gemini CLI
  • •Gemini Code Assist
  • •free tier
  • •rate limits
  • •AI coding tools

Related Articles

Explore more articles on similar topics to deepen your understanding of usage-based billing.

Gemini 3.8 Flash Pricing & Free Tier (October 2026)

Gemini 3.8 Flash is free on the Gemini API Free Tier and costs $0.75 input and $3.75 output per 1M tokens through Decemb...

9 min readRead more

OpenAI Codex Pricing & Usage Limits (October 2026): Every Plan

OpenAI Codex pricing as of October 2026: Free, Go $8, Plus $20, Pro from $100, Business $20 per user, and API key billin...

9 min readRead more

OpenAI API Usage Tiers 2026: Free, Build, Launch & Grow

OpenAI cut its paid API usage tiers from five to three on October 6, 2026: Build ($5 in credit purchases, $500/month), L...

8 min readRead more

Explore More Articles

Discover our complete collection of usage-based billing guides and implementation patterns.

View all articles