GPT-6 API Pricing (October 2026): Astra, Sol & Luna

GPT-6 API prices per 1M tokens as of October 2026: Astra $10/$50, GPT-6.1 Sol $2/$10, GPT-6 Sol $2/$10, Luna $0.10/$0.50. Cache writes, the 272K long-context surcharge, Batch, Flex, Fast and Ultrafast prices, release dates, and worked cost examples.

9 min read

GPT-6OpenAI API pricingGPT-6 AstraGPT-6.1 SolGPT-6 LunaAI cost

TL;DR (October 10, 2026): GPT-6 API prices per 1M tokens, standard processing, prompts up to 272K tokens: GPT-6 Astra $10 input / $50 output, GPT-6.1 Sol $2 / $10 (cached input $0.10), GPT-6 Sol $2 / $10 (cached input $0.20) and GPT-6 Luna $0.10 / $0.50. Prompts over 272K input tokens cost 2x on input and 1.5x on output for the whole request. Batch and Flex are half price, Fast is double, and Ultrafast (Astra and 6.1 Sol only) is 6x. For most coding and agent work, GPT-6.1 Sol is the default worth testing first: OpenAI positions it as near-Astra quality at one fifth of Astra's token price.

OpenAI shipped the GPT-6 family in three steps in September 2026, and the price list now has four GPT-6 models, a separate cache-write price, a long-context surcharge and five processing tiers. This page puts the numbers in one place, with the release dates, the context limits and worked cost examples, all taken from OpenAI's pricing page, model pages and API changelog as of October 2026.

GPT-6 API pricing table (standard processing)

ModelInputCached inputCache writesOutputInput over 272KOutput over 272K
gpt-6-astra$10.00$1.00$12.50$50.00$20.00$75.00
gpt-6.1-sol$2.00$0.10$2.50$10.00$4.00$15.00
gpt-6-sol$2.00$0.20$2.50$10.00$4.00$15.00
gpt-6-luna$0.10$0.01$0.125$0.50$0.20$0.75

USD per 1M tokens. "Short context" means 272K input tokens or fewer. Read on October 10, 2026 from OpenAI's API pricing page; the GPT-6 Sol row comes from its model page, because the pricing page's flagship table now lists only Astra, 6.1 Sol and Luna.

For comparison, the GPT-5.6 models are still on the list: GPT-5.6 Sol at $4 / $20 (a promotional price OpenAI says runs at least through November 21, 2026), GPT-5.6 Terra at $2 / $12 and GPT-5.6 Luna at $0.20 / $1.20. GPT-6 Sol costs half of GPT-5.6 Sol, and GPT-6 Luna costs half of GPT-5.6 Luna. Our earlier breakdown of GPT-5.6 pricing has the background on that generation.

Release dates and model limits

ModelAPI releaseContext windowMax outputKnowledge cutoff
GPT-6 AstraSeptember 3, 20261,050,000128,000April 30, 2026
GPT-6 SolSeptember 22, 20261,050,000128,000April 20, 2026
GPT-6 LunaSeptember 22, 20261,050,000128,000May 18, 2026
GPT-6.1 SolSeptember 29, 20261,050,000128,000April 30, 2026

OpenAI describes Astra as its most capable model, "built for the hardest end-to-end work". GPT-6.1 Sol is pitched as being for "complex coding and professional work at a lower cost than GPT-6 Astra", and its model page calls it "near-Astra performance for complex work". Sol and Luna are reasoning models with text and image input and text output. Luna was also the launch model for the Decisions API beta on October 6, 2026.

Processing tiers: Batch, Flex, Fast and Ultrafast

The same model has a different price depending on how you ask OpenAI to process it. Priority processing was renamed Fast mode on July 30, 2026.

ModelStandard (in / out)Batch and FlexFastUltrafast
GPT-6 Astra$10 / $50$5 / $25$20 / $100$60 / $300
GPT-6.1 Sol$2 / $10$1 / $5$4 / $20$12 / $60
GPT-6 Sol$2 / $10$1 / $5$4 / $20Not listed
GPT-6 Luna$0.10 / $0.50$0.05 / $0.25$0.20 / $1.00Not listed

Ultrafast arrived for GPT-6 Astra on September 29, 2026 (global processing and US data residency) and for GPT-6.1 Sol on October 8, 2026 (global, US and EU data residency). You request it with service_tier: "ultrafast" in the Responses API. It reduces the time between generated tokens, and it costs six times the standard rate, so it belongs on interactive paths where a person is waiting, not on background jobs.

Two surcharges apply on top: regional processing (data residency) endpoints add 10% for models released on or after March 5, 2026, which includes every GPT-6 model, and FedRAMP endpoints also add 10%.

The 272K long-context rule

Every GPT-6 model page carries the same note: prompts with more than 272K input tokens are priced at 2x the input and cache rates and 1.5x the output rate for the full request. The surcharge is not applied only to the tokens above 272K. Crossing the line reprices everything.

On GPT-6 Astra, a request with 270K input tokens and 5K output tokens costs $2.70 plus $0.25, or $2.95. Add 30K more input tokens to reach 300K and the same request costs $6.00 plus $0.375, or $6.38. Eleven percent more input more than doubles the bill. If you build agents that keep appending tool output to the context, trimming or compacting before the 272K mark is one of the largest savings available.

Caching: why 6.1 Sol is cheaper than 6 Sol

GPT-6.1 Sol and GPT-6 Sol share input and output prices, so the difference sits in caching. Cached input on 6.1 Sol is 5% of the uncached rate ($0.10), against 10% on GPT-6 Sol ($0.20). Cache writes are billed at 1.25x the uncached input rate on all four models, so the first request that creates a cache entry costs a little more than an uncached one.

For workloads with a long stable prefix (a system prompt, tool definitions, a repository map), that difference adds up. Our prompt caching cost comparison explains how to measure your real hit rate before counting on it.

Worked example: one workload, every model

Take 10,000 requests a day, each with 8,000 input tokens and 1,000 output tokens, under the 272K threshold. That is 80M input and 10M output tokens a day. At standard prices, without caching:

ModelPer dayPer 30 daysPer day with 75% of input cached
GPT-6 Astra$1,300$39,000$760
GPT-6.1 Sol$260$7,800$146
GPT-6 Sol$260$7,800$152
GPT-6 Luna$13$390$7.60
GPT-5.6 Sol$520$15,600not calculated
GPT-5.6 Terra$280$8,400not calculated

The cached column assumes 6,000 of the 8,000 input tokens are cache reads and ignores the one-off cache-write premium. Reasoning tokens are billed as output, so real output counts on reasoning models are often higher than the visible answer.

The ratio is the useful part: Astra costs five times as much as 6.1 Sol on the same token counts, and 100 times as much as Luna. A routing layer that sends only the hard tasks to Astra usually pays for itself quickly. If you compare against other vendors, remember that token counts differ between tokenizers, which we covered in why the token count is not the bill.

Migration notes that affect cost

  • Astra drops some parameters. It does not support the none reasoning effort, custom temperature or top_p, or logprobs. Tool calling requires the Responses API.
  • 6.1 Sol always reasons. Its reasoning effort starts at low; none and minimal are not supported, so you cannot use it as a no-reasoning model to save output tokens.
  • Sol and Luna support none. On Chat Completions they support function calling only with reasoning_effort set to none.
  • Rate limits come with your tier. On the Build tier, Astra and both Sol models get 5,000 RPM and 1M TPM, and Luna gets 2M TPM. See OpenAI API usage tiers for the full ladder.

Keeping the bill attributable

With four models, five processing tiers, a cache-write price and a long-context multiplier, the price of a single request now depends on six or more inputs. The OpenAI invoice tells you the total per model. It does not tell you which customer or feature produced it. If you resell or bundle GPT-6 usage, record model, service tier, cached tokens and input size with each request on your side. That is what UsageBox meters: the usage events your own application emits, which you can then reconcile against the OpenAI bill.

Sources

Read on October 10, 2026: the OpenAI API pricing page, the OpenAI API changelog (September 3, September 22, September 29, October 6 and October 8, 2026 entries), and the OpenAI model pages for GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna.

Key Topics

  • •GPT-6
  • •OpenAI API pricing
  • •GPT-6 Astra
  • •GPT-6.1 Sol
  • •GPT-6 Luna
  • •AI cost

Related Articles

Explore more articles on similar topics to deepen your understanding of usage-based billing.

Gemini 3.8 Flash Pricing & Free Tier (October 2026)

Gemini 3.8 Flash is free on the Gemini API Free Tier and costs $0.75 input and $3.75 output per 1M tokens through Decemb...

9 min readRead more

OpenAI API Usage Tiers 2026: Free, Build, Launch & Grow

OpenAI cut its paid API usage tiers from five to three on October 6, 2026: Build ($5 in credit purchases, $500/month), L...

8 min readRead more

The AI-Wrapper Margin: How to Find Out What That $29/Month Tool Actually Pays Per Token (2026)

A screenshot recipe went around this month for reverse-engineering what a $29/month AI tool actually pays: find the mode...

7 min readRead more

Explore More Articles

Discover our complete collection of usage-based billing guides and implementation patterns.

View all articles