OpenAI Cuts GPT-5.6 Sol API Output Pricing by One-Third in Three-Month Promotion
OpenAI cut GPT-5.6 Sol API output pricing to $20 per million tokens and input to $4 through at least November 21.
1. OpenAI lowers the cost of its flagship model
OpenAI reduced the API and credit pricing of GPT-5.6 Sol on August 21, 2026, introducing promotional rates that the company says will remain available at least through November 21. The largest change is the output-token price: standard short-context API output now costs $20 per million tokens, down from $30.
The standard input price fell from $5 to $4 per million tokens. Cached input dropped from $0.50 to $0.40, while cache writes—which are billed at 1.25 times the uncached input rate—fell from $6.25 to $5 per million tokens.
| Standard short-context usage | Previous price per 1M tokens | Promotional price | Reduction |
|---|---|---|---|
| Input | $5.00 | $4.00 | 20% |
| Cached input | $0.50 | $0.40 | 20% |
| Cache writes | $6.25 | $5.00 | 20% |
| Output | $30.00 | $20.00 | 33.3% |
OpenAI described the overall reduction as more than 20%, but the saving for an individual workload depends on its input-to-output ratio. A request consuming one million input tokens and producing one million output tokens would have cost $35 under the previous tariff. It now costs $24, a reduction of approximately 31.4%.
Output-heavy workloads receive the greatest percentage benefit. One million input tokens followed by three million output tokens now cost $64 instead of $95, a reduction of approximately 32.6%. That matters for coding agents, research systems and long-form generation workflows in which the model produces substantial reasoning, code or documentation.
The change affects billing rather than the model identifier. Developers can continue using gpt-5.6-sol; the gpt-5.6 alias still routes to Sol. OpenAI’s documentation continues to list a 1.05-million-token context window, a maximum output of 128,000 tokens, text and image input, function calling, structured outputs and tools including web search, file search, code execution, computer use and MCP.
2. Purchased credits become cheaper, but included usage does not expand
The reduction is also rolling out to eligible ChatGPT Work and Codex usage paid for with purchased credits. OpenAI’s credit-based rate card now charges 100 credits per million input tokens, 10 credits per million cached-input tokens and 500 credits per million output tokens for GPT-5.6 Sol.
Those rates apply to supported agentic activity, including local and cloud tasks, automations, delegated workers and eligible Codex workflows. OpenAI says a typical Codex task using Sol may consume between five and 30 credits, although the actual amount depends on token volume, reasoning, task complexity, automation and whether faster processing is selected.
The promotion does not increase included subscription allowances. OpenAI states that included plan usage, five-hour and weekly limits, and legacy credit rates remain unchanged. The ordinary ChatGPT credit rate for a Sol message also remains 10 credits, so the promotion should not be interpreted as a larger message quota for ChatGPT subscribers.
For standard ChatGPT Business seats, included limits are consumed before eligible work can draw from a workspace’s purchased credit pool. Usage-based Codex-only seats require workspace credits from the beginning. The current rate card also covers new and existing Enterprise, Edu, Health, Gov and ChatGPT for Teachers customers using the applicable flexible-pricing structure.
A small subset of Enterprise customers remains on a legacy Codex rate card until migration. Those workspaces do not automatically receive the new token-based credit calculation. Customers whose contracts specify usage-based billing in US dollars must use the tariff in their agreement rather than assuming the general credit table applies.
This distinction explains why OpenAI could lower “credit pricing” while saying that Pro, Plus and Business subscription usage remains unchanged. Purchased, token-metered usage can go further; an existing subscription’s included limits do not become larger.
3. Long context and faster processing still carry different rates
The headline $4 input and $20 output prices apply to Standard processing with no more than 272,000 input tokens. When a prompt exceeds that threshold, OpenAI charges the long-context rates for the entire request, not merely for the tokens above 272,000.
Under the promotional tariff, long-context Standard usage costs $8 per million input tokens, $0.80 per million cached-input tokens, $10 per million cache-write tokens and $30 per million output tokens. That represents a two-times multiplier for input and cached input and a 1.5-times multiplier for output.
Processing tier also changes the bill:
| GPT-5.6 Sol tier, short context | Input per 1M tokens | Cached input | Cache writes | Output |
|---|---|---|---|---|
| Standard | $4.00 | $0.40 | $5.00 | $20.00 |
| Batch | $2.00 | $0.20 | $2.50 | $10.00 |
| Flex | $2.00 | $0.20 | $2.50 | $10.00 |
| Fast mode | $8.00 | $0.80 | $10.00 | $40.00 |
Batch and Flex therefore remain half the Standard price for workloads that can accept their execution characteristics. Fast mode costs twice the promotional Standard rate. OpenAI renamed Priority Processing to Fast mode on July 30, but requests using service_tier: "priority" remain compatible.
Eligible regional-processing endpoints carry a further 10% data-residency uplift. OpenAI also warns that models supplied through Amazon Bedrock are billed by AWS and may not match the direct OpenAI API tariff.
For developers estimating savings, the processing tier, context length and cache behavior can be more important than the headline percentage. A long-context Fast-mode request, for example, is listed at $16 per million input tokens and $60 per million output tokens—four and three times the respective short-context Standard promotional rates.
4. Sol follows earlier Terra and Luna price cuts
The Sol promotion is OpenAI’s second GPT-5.6 pricing move in less than a month. On July 30, the company cut GPT-5.6 Terra pricing by 20% and Luna pricing by 80%. Sol remained at its launch price of $5 per million input tokens and $30 per million output tokens at that time, while OpenAI introduced Fast mode at twice the Standard price and up to 2.5 times the speed.
OpenAI attributed those earlier changes to improvements across its inference systems and production software. The company reported that Sol-assisted kernel optimization reduced its end-to-end serving cost by 20%, while experiments involving token generation improved efficiency by more than 15%. Those figures are company-reported engineering results, not an independent audit of OpenAI’s costs.
Reuters placed the new Sol discount in the context of competition from Anthropic and Chinese model providers. The practical competitive change is measurable without speculating about OpenAI’s motives: Sol is now cheaper than GPT-5.5’s published $5 input and $30 output tariff, while retaining OpenAI’s flagship position in the GPT-5.6 family.
The promotion does not eliminate the price differences among the three GPT-5.6 tiers. For a simplified workload using one million input and one million output tokens, Standard short-context pricing totals $24 on Sol, $14 on Terra and $1.40 on Luna. Developers still need workload-specific evaluations to decide when Sol’s higher capability justifies its cost.
OpenAI has committed only to making the promotional Sol pricing available at least through November 21. It has not said whether the rates will become permanent, be extended or return to the previous tariff. Cost forecasts extending beyond that date should therefore retain the former $5 input and $30 output prices as a fallback scenario.
Frequently Asked Questions
What are the new GPT-5.6 Sol API prices?
Standard short-context usage costs $4 per million input tokens, $0.40 per million cached-input tokens and $20 per million output tokens.
How long will the lower prices last?
OpenAI says the promotional pricing will be available at least through November 21, 2026. It has not announced what will replace it.
Do ChatGPT subscription limits increase?
No. Included plan usage, five-hour and weekly limits, and legacy credit rates remain unchanged. The reduction applies to API usage and eligible activity paid for with purchased credits.
Do API developers need to change model IDs?
No. The model remains gpt-5.6-sol, and the gpt-5.6 alias continues to route to Sol.
Does the $4 input rate apply to prompts over 272,000 tokens?
No. Requests exceeding 272,000 input tokens are charged long-context rates for the full request, including $8 per million input tokens and $30 per million output tokens under Standard processing.
Sources
Share