OpenAI API Usage Tiers 2026: Free, Build, Launch & Grow

OpenAI cut its paid API usage tiers from five to three on October 6, 2026: Build ($5 in credit purchases, $500/month), Launch ($100, $5,000/month) and Grow ($500, $200,000/month). GPT-6 rate limits per tier, spend limits vs usage limits, the slow_down 429, and rate limit headers.

8 min read

OpenAI APIusage tiersrate limitsspend limitsGPT-6

TL;DR (October 10, 2026): On October 6, 2026 OpenAI cut its API usage tiers from five paid tiers to three. The ladder is now Free ($100/month usage limit), Build ($5 in total credit purchases, $500/month), Launch ($100 in purchases, $5,000/month) and Grow ($500 in purchases, $200,000/month). Upgrades are automatic once your all-time credit purchases cross the threshold, with no waiting period listed. For the GPT-6 family, Build gets 5,000 RPM, Launch 10,000 RPM and Grow 15,000 RPM (30,000 for GPT-6 Luna). The tier's monthly usage limit is a separate ceiling from the spend limits you set yourself, and each one fails with its own error code.

If you set up an OpenAI API account before October, you were placed in one of five usage tiers. That structure is gone. The rate limits guide and the API changelog entry for October 6, 2026 describe a simpler ladder: three named paid tiers, each unlocked by how much you have paid in credits in total. This page covers the new tiers, the per-model limits that come with them, and the three different ceilings that can stop your traffic.

OpenAI API usage tiers as of October 2026

TierHow you qualifyMonthly usage limit
FreeUser must be in an allowed geography$100 / month
Build$5 in total credit purchases$500 / month
Launch$100 in total credit purchases$5,000 / month
Grow$500 in total credit purchases$200,000 / month

Source: OpenAI API rate limits guide, usage tiers section, read October 10, 2026. Qualification is cumulative: it counts every credit purchase on the organization, not one month's spend.

Two details matter more than the table suggests. First, the qualification is about money paid in, not money used. Buying $100 of credits moves an organization to Launch even if it has spent only a few dollars of them. Second, the guide lists no time condition for any tier, so on paper an organization can reach Grow on its first day by buying $500 of credits.

What changed on October 6, 2026

The changelog entry is short: "Simplified API usage tiers from five to three: Build, Launch, and Grow. Organizations automatically upgrade as total credit purchases reach tier minimums." In practice that means:

  • Fewer steps. There are three paid rungs instead of five. The top rung's monthly usage limit is $200,000.
  • Purchase thresholds only. Every paid tier now qualifies on one thing: total credit purchases. If your internal runbook still refers to numbered tiers, it describes a structure the docs no longer use.
  • Limits live in the dashboard. The guide sends you to Settings, then Organization, then Limits, to see per-model rate limits for your tier, and there is an Upgrade tier control in the Usage Tiers section.

Rate limits per model at each tier

Each model page lists its own limits by tier. For the current GPT-6 models the numbers are:

ModelBuild (RPM / TPM)Launch (RPM / TPM)Grow (RPM / TPM)
gpt-6-astra5,000 / 1,000,00010,000 / 4,000,00015,000 / 40,000,000
gpt-6.1-sol5,000 / 1,000,00010,000 / 4,000,00015,000 / 40,000,000
gpt-6-sol5,000 / 1,000,00010,000 / 4,000,00015,000 / 40,000,000
gpt-6-luna5,000 / 2,000,00010,000 / 10,000,00030,000 / 180,000,000

Batch work has its own ceiling, the batch queue limit, which counts the input tokens of all pending batch jobs for a model. On the model pages that publish it, GPT-6 Astra and GPT-6 Sol allow 3,000,000 queued tokens on Build, 200,000,000 on Launch and 15,000,000,000 on Grow. GPT-6 Luna allows 20,000,000, 1,000,000,000 and 15,000,000,000. Once a batch completes, its tokens stop counting.

The jump from Build to Launch is the one most small teams feel. On GPT-6 Sol the token limit goes from 1M to 4M TPM, and that costs a cumulative $95 more in credit purchases. If a nightly job keeps hitting 429s on Build, buying credits you will use anyway is often the cheapest fix.

The Free tier is not on that ladder for GPT-6. Every GPT-6 model page lists Free as "Not supported", so a free organization needs at least the $5 Build purchase before it can call Astra, Sol or Luna. For older models, check the Limits page in the dashboard for what a free organization actually gets.

How the limits are counted

OpenAI measures limits in requests per minute (RPM), requests per day (RPD), tokens per minute (TPM), tokens per day (TPD) and images per minute (IPM), plus audio minutes per minute for some streaming audio models. You hit whichever one runs out first. The guide's own example: 20 requests of 100 tokens each fill a 20 RPM limit even though they use a tiny fraction of a 150k TPM limit.

  • Organization and project, not user. Limits apply at the organization level and at the project level. Two developers sharing a project share its limits.
  • Shared limits. Some model families share one budget. If the Limits page lists models under a "shared limit" with 3.5M TPM, calls to any of them count toward the same 3.5M.
  • Long context is separate. Long-context models such as GPT-5.5 have their own limit for long-context requests, shown in the dashboard.
  • Failed requests count. The guide says unsuccessful requests contribute to your per-minute limit, so retrying in a tight loop makes things worse. We cover the cost side of that in does a 429 count against your rate limit.
  • Vector stores. File ingestion into one vector store is capped at 300 requests per minute across the files and file_batches endpoints.

Four ceilings that can stop your traffic

Rate limits are only one of four things that can stop traffic. The other three are described in the spend limits guide, each returns its own error code, and mixing them up is the most common reason a team "raises the limit" and still gets errors.

CeilingWho sets itWhat you seeHow to clear it
Rate limit (RPM, TPM and so on)OpenAI, by tier and model429, with rate limit headersBack off and retry, or move up a tier
Monthly usage limitOpenAI, by tierorganization_usage_limit_exceededRequest a higher approved usage limit
Hard spend limit (org)You429 with organization_spend_limit_exceededRaise or remove it, or wait for the monthly reset
Hard spend limit (project)You429 with project_spend_limit_exceededSame, at project level
Prepaid balanceYour credit balancecredit_balance_exhaustedAdd credits

Spend alerts sit next to these: they send a notification and let traffic continue. OpenAI is explicit that hard-limit enforcement "is not instantaneous", so recorded spend can slightly exceed the amount you set. Treat a hard limit as a backstop, not as an exact budget. If you need a stop that holds per customer rather than per project, you have to enforce it in your own code, which is the pattern in hard AI spend caps and kill switches.

Ramp-rate errors: the 429 you get while under your limits

Since September 2, 2026 the API separates two conditions that used to look alike. Traffic that grows too quickly returns 429 with the code slow_down. A temporarily overloaded model returns 503 with server_is_overloaded. Both can carry a Retry-After header.

The important line in the guide: a slow_down error "can occur even when your traffic is within its requests-per-minute and tokens-per-minute limits". It reacts to how fast traffic grew, not to how much you sent. OpenAI's rule of thumb is that once you reach 1 million input tokens per minute, you should increase traffic by no more than 50% every 15 minutes. A backfill that jumps from zero to full speed can trip this on any tier.

Reading your limits from response headers

You do not have to wait for an error to know where you stand. Responses can include:

  • x-ratelimit-limit-requests and x-ratelimit-limit-tokens: the ceiling
  • x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens: what is left
  • x-ratelimit-reset-requests and x-ratelimit-reset-tokens: time until reset
  • x-ratelimit-limit-project-tokens, x-ratelimit-remaining-project-tokens, x-ratelimit-reset-project-tokens: present when a project-scoped token limit applies
  • Retry-After: the minimum number of seconds to wait, when present

Logging the remaining-tokens header next to each request costs one field and gives you a history of headroom, which is far more useful than a pile of 429s when you are deciding whether to move up a tier.

What the monthly usage limit buys in tokens

The dollar ceilings are easier to plan around once you turn them into tokens. Using GPT-6.1 Sol's standard price of $2 per million input tokens and $10 per million output tokens (prompts up to 272K tokens), the Build tier's $500 a month covers, for example, 100M input tokens plus 30M output tokens ($200 plus $300). Launch's $5,000 covers ten times that. On GPT-6 Luna at $0.10 and $0.50, the same $500 covers 2.5 billion input tokens plus 500 million output tokens. The full price list is in our GPT-6 API pricing breakdown, and the previous generation is covered in GPT-5.6 pricing.

Practical checklist

  1. Check your tier today. Settings, Organization, Limits. If you were on an old numbered tier, confirm where you landed.
  2. Separate projects by workload. Project-level spend limits and rate limits are the only isolation OpenAI gives you. A batch job and a customer-facing chat should not share one project.
  3. Handle 429 and 503 separately. Inspect error.code: slow_down means ramp more gently, the spend codes mean a person has to act, and quota errors should never be retried.
  4. Set an alert below every hard limit. Alerts stay active when a hard limit exists, so you hear about it before traffic stops.
  5. Meter per customer yourself. OpenAI's limits stop at the project. If you resell model access, you need your own per-customer usage record, which is the job UsageBox does: it records the usage events your application emits, it does not read or change OpenAI's limits.

Sources

All figures were read on October 10, 2026 from OpenAI's own documentation: the OpenAI API rate limits guide, the OpenAI spend limits guide, the OpenAI API changelog (entries for September 2 and October 6, 2026), and the GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna model pages.

Key Topics

  • •OpenAI API
  • •usage tiers
  • •rate limits
  • •spend limits
  • •GPT-6

Related Articles

Explore more articles on similar topics to deepen your understanding of usage-based billing.

Gemini API Spend Caps & Tiers (2026): The $250 Hard Stop Nobody Read About

Since April 1, 2026 every Gemini API billing account has a mandatory monthly spend cap by tier (~$250 Tier 1, ~$2,000 Ti...

10 min readRead more

Gemini 3.8 Flash Pricing & Free Tier (October 2026)

Gemini 3.8 Flash is free on the Gemini API Free Tier and costs $0.75 input and $3.75 output per 1M tokens through Decemb...

9 min readRead more

Gemini CLI Free Tier Limits 2026: 1,000 Requests a Day

Gemini CLI limits by sign-in method as of October 2026: 1,000 requests a day with a Google account, 250 a day (Flash onl...

8 min readRead more

Explore More Articles

Discover our complete collection of usage-based billing guides and implementation patterns.

View all articles