To cite an historic Jedi Grasp "Begun, the AI price wars have!"
OpenAI is sharply lowering the costs of two fashions in its GPT-5.6 frontier sequence, slicing GPT-5.6 Luna, the smallest and quickest mannequin within the sequence, by 80% and GPT-5.6 Terra, the mid-tier mannequin, by 20%, whereas including a premium Quick mode for its flagship GPT-5.6 Sol mannequin.
The cuts place Luna a lot nearer to the lowest-cost industrial fashions out there and arrive just some days after Anthropic released its highly performant Claude Opus 5 on the identical value as Opus 4.8, and Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two rival fashions constructed round decrease inference prices, sooner execution and extra environment friendly agent workloads.
OpenAI is efficiently undercutting Google's value per intelligence and making an attempt to sway Anthropic customers, who might not thoughts paying extra, with a pace enhance.
OpenAI says Luna will now price $0.20 per million enter tokens and $1.20 per million output tokens, for a mixed input-plus-output value of $1.40 per million tokens.
Terra will price $2 per million enter tokens and $12 per million output tokens, for a mixed value of $14.
Pricing for Sol Commonplace stays unchanged at $5 per million enter tokens and $30 per million output tokens. OpenAI can also be including Sol Quick mode at twice the Commonplace value: $10 per million enter tokens and $60 per million output tokens.
The corporate says Quick mode delivers as much as 2.5 occasions the throughput with out altering the mannequin’s underlying intelligence.
OpenAI co-founder and CEO Sam Altman took to X to announce the adjustments as "main value cuts at this time."
VentureBeat Frontier AI mannequin API pricing comparability
|
Mannequin |
Enter ($/1M) |
Output ($/1M) |
Complete ($/1M) |
Supply |
|
MiMo-V2.5 Flash |
$0.10 |
$0.30 |
$0.40 |
|
|
deepseek-v4-flash |
$0.14 |
$0.28 |
$0.42 |
|
|
deepseek-v4-pro |
$0.435 |
$0.87 |
$1.305 |
|
|
GPT-5.6 Luna |
$0.20 |
$1.20 |
$1.40 |
|
|
MiniMax-M3 |
$0.30 |
$1.20 |
$1.50 |
|
|
LongCat-2.0 — limited-time promo |
$0.30 |
$1.20 |
$1.50 |
|
|
Gemini 3.1 Flash-Lite |
$0.25 |
$1.50 |
$1.75 |
|
|
Qwen3.7-Plus |
$0.40 |
$1.60 |
$2.00 |
|
|
MiMo-V2.5 |
$0.40 |
$2.00 |
$2.40 |
|
|
Gemini 3.5 Flash-Lite |
$0.30 |
$2.50 |
$2.80 |
|
|
LongCat-2.0 — commonplace |
$0.75 |
$2.95 |
$3.70 |
|
|
MiMo-V2.5 Professional (≤256K) |
$1.00 |
$3.00 |
$4.00 |
|
|
GLM-5.2 |
$1.40 |
$4.40 |
$5.80 |
|
|
Grok 4.5 |
$2.00 |
$6.00 |
$8.00 |
|
|
MiMo-V2.5 Professional (>256K) |
$2.00 |
$6.00 |
$8.00 |
|
|
Gemini 3.6 Flash |
$1.50 |
$7.50 |
$9.00 |
|
|
Qwen3.7-Max |
$2.50 |
$7.50 |
$10.00 |
|
|
Gemini 3.5 Flash |
$1.50 |
$9.00 |
$10.50 |
|
|
Gemini 3.1 Professional Preview (≤200K) |
$2.00 |
$12.00 |
$14.00 |
|
|
GPT-5.6 Terra |
$2.00 |
$12.00 |
$14.00 |
|
|
GPT-5.4 |
$2.50 |
$15.00 |
$17.50 |
|
|
Kimi K3 |
$3.00 |
$15.00 |
$18.00 |
|
|
Gemini 3.1 Professional Preview (>200K) |
$4.00 |
$18.00 |
$22.00 |
|
|
Claude Opus 5 |
$5.00 |
$25.00 |
$30.00 |
|
|
GPT-5.5 |
$5.00 |
$30.00 |
$35.00 |
|
|
GPT-5.5 Immediate (chat-latest) |
$5.00 |
$30.00 |
$35.00 |
|
|
Sakana Fugu Extremely (≤272K) |
$5.00 |
$30.00 |
$35.00 |
|
|
GPT-5.6 Sol — Commonplace mode |
$5.00 |
$30.00 |
$35.00 |
|
|
Claude Fable 5 / Claude Mythos 5 |
$10.00 |
$50.00 |
$60.00 |
|
|
GPT-5.6 Sol — Quick mode |
$10.00 |
$60.00 |
$70.00 |
Pricing is proven per a million tokens. Complete price is calculated as enter value plus output value. Cached-input pricing is excluded to maintain the comparability constant throughout suppliers.
OpenAI strikes Luna into the low-cost tier
Probably the most consequential change is the Luna value lower.
When OpenAI launched the GPT-5.6 sequence, Luna was priced at $1 per million enter tokens and $6 per million output tokens, for a mixed complete of $7. The brand new pricing reduces that mixed determine to $1.40.
That locations Luna under Google’s Gemini 3.5 Flash-Lite, which prices a mixed $2.80 per million enter and output tokens, and much under Gemini 3.6 Flash at $9. Luna additionally now prices lower than OpenAI’s personal GPT-5.4 and Terra fashions by a large margin.
It’s not the most cost effective mannequin within the broader market. Xiaomi’s MiMo-V2.5 Flash, DeepSeek’s flash mannequin and several other different APIs stay cheaper on a pure token foundation. However the discount brings an OpenAI frontier-series mannequin into direct competitors with the market’s low-cost inference tier.
OpenAI says the GPT-5.6 sequence represents its frontier mannequin household, with Sol positioned on the prime of the lineup, Terra as the center tier and Luna because the smallest and quickest possibility.
The lineup was initially released in late June 2026 by means of a restricted rollout by U.S. authorities request, earlier than broader entry, with every mannequin supposed to supply a distinct tradeoff amongst intelligence, latency and value.
Sol is aimed on the most complicated reasoning-heavy and agentic workloads, together with superior coding, multi-step planning and tool-using methods, whereas Terra is designed for common manufacturing use the place a steadiness of functionality and effectivity is required. Luna is positioned for high-throughput, low-latency duties reminiscent of summarization, classification, routing, and light-weight real-time assistants the place price per request is the first constraint.
Terra drops to match Google’s Gemini 3.1 Professional pricing
Terra’s 20% discount strikes its mixed value from $17.50 to $14 per million tokens.
At that stage, Terra now matches Google’s Gemini 3.1 Professional Preview pricing for context home windows of 200,000 tokens or much less.
It additionally undercuts OpenAI’s GPT-5.4, which stays priced at $2.50 per million enter tokens and $15 per million output tokens, providing the identical intelligence for about 1/thirteenth the associated fee, as Krea AI's Nic Dunz noted on X:
The adjustment creates a wider separation between OpenAI’s three GPT-5.6 tiers. Luna prices one-tenth as a lot as Terra on a easy mixed input-plus-output foundation, whereas Terra prices 60% lower than Sol Commonplace.
Sol Quick strikes in the wrong way. At a mixed $70 per million tokens, it’s the costliest mannequin configuration within the comparability under, reflecting OpenAI’s resolution to cost a premium for latency-sensitive workloads fairly than decrease Sol’s base value.
Cuts observe Google’s low-cost Gemini releases and Anthropic's Claude Opus 5
OpenAI’s pricing adjustments come solely a few week and a half after Google introduced its own low-cost Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
Google priced Gemini 3.6 Flash at $1.50 per million enter tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite prices $0.30 per million enter tokens and $2.50 per million output tokens.
Google framed each fashions across the economics of agent deployment, arguing that decrease token utilization, fewer reasoning steps and decreased instrument calls may decrease the whole price of long-running software program engineering and knowledge-work duties.
Gemini 3.6 Flash reportedly makes use of 17% fewer output tokens than Gemini 3.5 Flash on the Synthetic Evaluation Index, with financial savings reaching as excessive as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned because the quickest mannequin in Google’s 3.5 sequence.
Nevertheless, OpenAI's fashions are extra performant than Google's, in line with third get together evaluation outfits like Artificial Analysis, with even the Luna mannequin outperforming Gemini 3.6 Flash and the older Gemini 3.1 Professional mannequin, making the cost-per intelligence way more favorable to OpenAI.
As AI coding startup Cognition noted on X, GPT-5.6 now "sits on the pareto curve of value/efficiency effectivity," posting an animation of the GPT-5.6 sequence shifting left on a chart representing intelligence on the y axis and value on the x, displaying that the fashions now provide among the many most superior intelligence for lowest price in the marketplace.
And but, rival Anthropic's Claude Opus 5 stays about as performant as GPT-5.6 Sol, but is 6% cheaper.
The mannequin prices $5 per million enter tokens and $25 per million output tokens—the identical charges as Opus 4.8—however Anthropic says it delivers almost all of the intelligence of its dearer Fable 5 mannequin at roughly half the associated fee.
Not like OpenAI’s Luna and Terra adjustments, Anthropic didn’t cut back the Opus API sticker value. As a substitute, it successfully lowered the worth per unit of functionality by changing Opus 4.8 with a extra succesful mannequin on the identical $30 mixed input-and-output charge. Anthropic additionally added an adjustable effort setting that permits builders to commerce reasoning depth for pace and token financial savings.
That distinction issues for enterprise consumers. OpenAI is instantly slicing per-token charges, Google is pairing decrease costs with reductions in token use and power calls, and Anthropic is emphasizing stronger activity efficiency at an unchanged value. All three approaches goal the identical operational metric: the whole price of finishing manufacturing work, fairly than the marketed price of a person token alone.
The timing highlights how shortly pricing has develop into a aggressive lever amongst frontier mannequin suppliers. OpenAI’s response doesn’t introduce a brand new mannequin era. As a substitute, it adjustments the economics of deploying fashions that have been launched solely just lately.
The market shifts from mannequin entry to mannequin economics
The cuts point out that entry to frontier-level functionality is now not the one level of competitors. The following query for enterprises is how cheaply and predictably these fashions can run in manufacturing.
OpenAI remains to be not the lowest-priced supplier on a pure token foundation. However Luna’s 80% discount materially adjustments its place, shifting it from the center of the market right into a pricing tier populated by smaller fashions from Google, Xiaomi, DeepSeek, MiniMax and different distributors.
That issues most for high-volume functions, the place comparatively small variations in token pricing can compound throughout coding brokers, doc methods, inside search instruments and automatic workflows.
OpenAI’s newest transfer due to this fact seems to be much less like a routine adjustment and extra like a repositioning of the GPT-5.6 sequence. Sol stays the premium possibility, Terra strikes nearer to competing pro-tier methods, and Luna turns into the corporate’s direct reply to the business’s rising low-cost mannequin phase.
