TL;DR (October 10, 2026): GPT-6 API prices per 1M tokens, standard processing, prompts up to 272K tokens: GPT-6 Astra $10 input / $50 output, GPT-6.1 Sol $2 / $10 (cached input $0.10), GPT-6 Sol $2 / $10 (cached input $0.20) and GPT-6 Luna $0.10 / $0.50. Prompts over 272K input tokens cost 2x on input and 1.5x on output for the whole request. Batch and Flex are half price, Fast is double, and Ultrafast (Astra and 6.1 Sol only) is 6x. For most coding and agent work, GPT-6.1 Sol is the default worth testing first: OpenAI positions it as near-Astra quality at one fifth of Astra's token price.
OpenAI shipped the GPT-6 family in three steps in September 2026, and the price list now has four GPT-6 models, a separate cache-write price, a long-context surcharge and five processing tiers. This page puts the numbers in one place, with the release dates, the context limits and worked cost examples, all taken from OpenAI's pricing page, model pages and API changelog as of October 2026.
GPT-6 API pricing table (standard processing)
| Model | Input | Cached input | Cache writes | Output | Input over 272K | Output over 272K |
|---|---|---|---|---|---|---|
| gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 | $20.00 | $75.00 |
| gpt-6.1-sol | $2.00 | $0.10 | $2.50 | $10.00 | $4.00 | $15.00 |
| gpt-6-sol | $2.00 | $0.20 | $2.50 | $10.00 | $4.00 | $15.00 |
| gpt-6-luna | $0.10 | $0.01 | $0.125 | $0.50 | $0.20 | $0.75 |
USD per 1M tokens. "Short context" means 272K input tokens or fewer. Read on October 10, 2026 from OpenAI's API pricing page; the GPT-6 Sol row comes from its model page, because the pricing page's flagship table now lists only Astra, 6.1 Sol and Luna.
For comparison, the GPT-5.6 models are still on the list: GPT-5.6 Sol at $4 / $20 (a promotional price OpenAI says runs at least through November 21, 2026), GPT-5.6 Terra at $2 / $12 and GPT-5.6 Luna at $0.20 / $1.20. GPT-6 Sol costs half of GPT-5.6 Sol, and GPT-6 Luna costs half of GPT-5.6 Luna. Our earlier breakdown of GPT-5.6 pricing has the background on that generation.
Release dates and model limits
| Model | API release | Context window | Max output | Knowledge cutoff |
|---|---|---|---|---|
| GPT-6 Astra | September 3, 2026 | 1,050,000 | 128,000 | April 30, 2026 |
| GPT-6 Sol | September 22, 2026 | 1,050,000 | 128,000 | April 20, 2026 |
| GPT-6 Luna | September 22, 2026 | 1,050,000 | 128,000 | May 18, 2026 |
| GPT-6.1 Sol | September 29, 2026 | 1,050,000 | 128,000 | April 30, 2026 |
OpenAI describes Astra as its most capable model, "built for the hardest end-to-end work". GPT-6.1 Sol is pitched as being for "complex coding and professional work at a lower cost than GPT-6 Astra", and its model page calls it "near-Astra performance for complex work". Sol and Luna are reasoning models with text and image input and text output. Luna was also the launch model for the Decisions API beta on October 6, 2026.
Processing tiers: Batch, Flex, Fast and Ultrafast
The same model has a different price depending on how you ask OpenAI to process it. Priority processing was renamed Fast mode on July 30, 2026.
| Model | Standard (in / out) | Batch and Flex | Fast | Ultrafast |
|---|---|---|---|---|
| GPT-6 Astra | $10 / $50 | $5 / $25 | $20 / $100 | $60 / $300 |
| GPT-6.1 Sol | $2 / $10 | $1 / $5 | $4 / $20 | $12 / $60 |
| GPT-6 Sol | $2 / $10 | $1 / $5 | $4 / $20 | Not listed |
| GPT-6 Luna | $0.10 / $0.50 | $0.05 / $0.25 | $0.20 / $1.00 | Not listed |
Ultrafast arrived for GPT-6 Astra on September 29, 2026 (global processing and US data residency) and for GPT-6.1 Sol on October 8, 2026 (global, US and EU data residency). You request it with service_tier: "ultrafast" in the Responses API. It reduces the time between generated tokens, and it costs six times the standard rate, so it belongs on interactive paths where a person is waiting, not on background jobs.
Two surcharges apply on top: regional processing (data residency) endpoints add 10% for models released on or after March 5, 2026, which includes every GPT-6 model, and FedRAMP endpoints also add 10%.
The 272K long-context rule
Every GPT-6 model page carries the same note: prompts with more than 272K input tokens are priced at 2x the input and cache rates and 1.5x the output rate for the full request. The surcharge is not applied only to the tokens above 272K. Crossing the line reprices everything.
On GPT-6 Astra, a request with 270K input tokens and 5K output tokens costs $2.70 plus $0.25, or $2.95. Add 30K more input tokens to reach 300K and the same request costs $6.00 plus $0.375, or $6.38. Eleven percent more input more than doubles the bill. If you build agents that keep appending tool output to the context, trimming or compacting before the 272K mark is one of the largest savings available.
Caching: why 6.1 Sol is cheaper than 6 Sol
GPT-6.1 Sol and GPT-6 Sol share input and output prices, so the difference sits in caching. Cached input on 6.1 Sol is 5% of the uncached rate ($0.10), against 10% on GPT-6 Sol ($0.20). Cache writes are billed at 1.25x the uncached input rate on all four models, so the first request that creates a cache entry costs a little more than an uncached one.
For workloads with a long stable prefix (a system prompt, tool definitions, a repository map), that difference adds up. Our prompt caching cost comparison explains how to measure your real hit rate before counting on it.
Worked example: one workload, every model
Take 10,000 requests a day, each with 8,000 input tokens and 1,000 output tokens, under the 272K threshold. That is 80M input and 10M output tokens a day. At standard prices, without caching:
| Model | Per day | Per 30 days | Per day with 75% of input cached |
|---|---|---|---|
| GPT-6 Astra | $1,300 | $39,000 | $760 |
| GPT-6.1 Sol | $260 | $7,800 | $146 |
| GPT-6 Sol | $260 | $7,800 | $152 |
| GPT-6 Luna | $13 | $390 | $7.60 |
| GPT-5.6 Sol | $520 | $15,600 | not calculated |
| GPT-5.6 Terra | $280 | $8,400 | not calculated |
The cached column assumes 6,000 of the 8,000 input tokens are cache reads and ignores the one-off cache-write premium. Reasoning tokens are billed as output, so real output counts on reasoning models are often higher than the visible answer.
The ratio is the useful part: Astra costs five times as much as 6.1 Sol on the same token counts, and 100 times as much as Luna. A routing layer that sends only the hard tasks to Astra usually pays for itself quickly. If you compare against other vendors, remember that token counts differ between tokenizers, which we covered in why the token count is not the bill.
Migration notes that affect cost
- Astra drops some parameters. It does not support the
nonereasoning effort, customtemperatureortop_p, orlogprobs. Tool calling requires the Responses API. - 6.1 Sol always reasons. Its reasoning effort starts at
low;noneandminimalare not supported, so you cannot use it as a no-reasoning model to save output tokens. - Sol and Luna support
none. On Chat Completions they support function calling only withreasoning_effortset tonone. - Rate limits come with your tier. On the Build tier, Astra and both Sol models get 5,000 RPM and 1M TPM, and Luna gets 2M TPM. See OpenAI API usage tiers for the full ladder.
Keeping the bill attributable
With four models, five processing tiers, a cache-write price and a long-context multiplier, the price of a single request now depends on six or more inputs. The OpenAI invoice tells you the total per model. It does not tell you which customer or feature produced it. If you resell or bundle GPT-6 usage, record model, service tier, cached tokens and input size with each request on your side. That is what UsageBox meters: the usage events your own application emits, which you can then reconcile against the OpenAI bill.
Sources
Read on October 10, 2026: the OpenAI API pricing page, the OpenAI API changelog (September 3, September 22, September 29, October 6 and October 8, 2026 entries), and the OpenAI model pages for GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna.