OpenAI simply made its least expensive frontier mannequin dramatically extra inexpensive — and the timing is something however unintended. The corporate slashed costs on two fashions in its GPT-5.6 sequence, slicing GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, whereas concurrently launching a premium Quick mode for its flagship GPT-5.6 Sol. CEO Sam Altman introduced the adjustments publicly on X, calling them “main value cuts at present.” The strikes land roughly three weeks after the GPT-5.6 sequence first grew to become broadly accessible — and simply days after rivals Google and Anthropic every made their very own cost-efficiency performs.
Key takeaways
- GPT-5.6 Luna’s mixed token value dropped 80% to $1.40 per million tokens, undercutting Google’s Gemini 3.5 Flash-Lite ($2.80) and Gemini 3.6 Flash ($9).
- GPT-5.6 Terra fell 20% from $17.50 to $14 per million tokens mixed, matching Google’s Gemini 3.1 Professional Preview for context home windows as much as 200,000 tokens.
- Sol Quick mode provides as much as 2.5x throughput at $70 per million tokens mixed — double the Normal charge — for latency-sensitive manufacturing workloads.
- Anthropic’s Claude Opus 5 stays priced at $30 mixed per million tokens, delivering near-Fable 5 efficiency on the similar charge as Opus 4.8.
- OpenAI’s Luna now competes instantly within the low-cost mannequin tier alongside choices from Google, Xiaomi, DeepSeek, and MiniMax.
OpenAI Broadcasts Main Value Cuts Throughout the GPT-5.6 Collection
The numbers inform a stark story. Luna, the smallest and quickest mannequin within the GPT-5.6 lineup, beforehand carried a mixed input-and-output value of $7 per million tokens. After the minimize, that determine collapses to $0.20 per million enter tokens and $1.20 per million output tokens — a mixed $1.40. For top-volume functions processing hundreds of thousands of requests every day, that hole is gigantic.
80% Value Discount for Luna
At $1.40 mixed per million tokens, Luna now sits beneath Google’s Gemini 3.5 Flash-Lite at $2.80, and much beneath Gemini 3.6 Flash at $9. It additionally undercuts OpenAI’s personal GPT-5.4. Luna doesn’t declare absolutely the lowest token value available in the market — fashions from Xiaomi, DeepSeek, and MiniMax nonetheless undercut it on uncooked price — nevertheless it marks the primary time an OpenAI frontier-series mannequin has entered that aggressive pricing band.
That repositioning issues past the numbers. Luna targets high-throughput, low-latency workloads — summarization, classification, routing, light-weight real-time assistants — the place the price per particular person request compounds at scale. Shifting into this tier means OpenAI is now competing instantly for workload classes it beforehand ceded to smaller, cheaper fashions.
20% Value Discount for Terra
Terra’s minimize is extra modest however strategically pointed. The mixed token value dropped from $17.50 to $14 per million tokens, matching Google’s Gemini 3.1 Professional Preview for context home windows of 200,000 tokens or much less. As Krea AI’s Nic Dunz famous on X, Terra additionally now undercuts OpenAI’s personal GPT-5.4, which stays priced at $2.50 per million enter and $15 per million output tokens — making Terra the higher worth at roughly one-thirteenth the price on a per-intelligence foundation. Terra is designed for common manufacturing deployments the place functionality and effectivity should be balanced, not maximized in a single route.
Introduction of Sol Quick Premium Mode
Sol strikes in the other way. Normal pricing stays at $5 per million enter tokens and $30 per million output tokens. The brand new Sol Quick mode expenses $10 per million enter tokens and $60 per million output tokens — a mixed $70 — delivering as much as 2.5 instances the throughput with out altering the underlying mannequin’s intelligence. Quite than reducing the value of its most succesful tier, OpenAI is charging a premium for latency benefits, signaling that for complicated reasoning and agentic workloads, velocity has its personal market.
Aggressive Positioning of OpenAI’s GPT-5.6 Fashions
The three GPT-5.6 tiers now map onto clearly distinct market segments. Luna competes within the low-cost inference market. Terra targets the mid-market professional tier. Sol anchors the frontier reasoning class.
Luna Competes within the Low-Price AI Section
The 80% Luna minimize transforms OpenAI’s aggressive footprint. Beforehand, the GPT-5.6 sequence was largely a premium providing. Now one tier sits inside the identical pricing bracket as fashions from Google, Xiaomi, DeepSeek, and MiniMax. Based on third-party evaluation from Synthetic Evaluation, Luna outperforms Gemini 3.6 Flash and Gemini 3.1 Professional in intelligence benchmarks — which means Luna’s cost-per-intelligence ratio has shifted materially in OpenAI’s favor. As AI startup Cognition famous on X, GPT-5.6 now “sits on the pareto curve of value/efficiency effectivity,” providing among the many most favorable intelligence-to-cost ratios in the marketplace.
Terra Matches Google’s Gemini 3.1 Professional Pricing
Terra’s $14 mixed value level creates a direct match with Gemini 3.1 Professional Preview for mid-range context workloads. The broader hole it creates inside OpenAI’s personal lineup is notable too: Luna now prices one-tenth of Terra on a easy combined-token foundation, whereas Terra prices 60% lower than Sol Normal. The three tiers are not intently spaced — they symbolize genuinely completely different price-performance trade-offs.
Sol Targets Advanced Reasoning Workloads
Sol’s positioning has not modified. It stays the mannequin for superior coding, multi-step planning, and tool-using agentic methods — workloads the place reasoning depth justifies larger per-token prices. The Quick mode addition extends Sol’s attraction to enterprises that want frontier intelligence however can’t take in the latency of Normal throughput. At $70 mixed per million tokens, Sol Quick is the most costly configuration within the lineup, a deliberate premium for time-critical deployments.
Market and Business Implications of the Value Cuts
OpenAI’s timing displays strain from a number of instructions concurrently. Based on reporting by CNBC, enterprises have grown more and more cost-sensitive, scrutinizing AI payments which have at instances reached billions of {dollars}. The period of limitless AI utilization with out monitoring prices — what some described as “tokenmaxxing” — has given strategy to a sharper deal with return on funding. That shift is forcing all frontier mannequin suppliers to rethink what they cost and why.
Shift from Mannequin Entry to Manufacturing Economics
What OpenAI described in its launch as a deal with “advancing each functionality and effectivity so every era of intelligence can accomplish extra work at a decrease price” displays one thing broader than a routine pricing adjustment. The competitors amongst frontier suppliers has moved past which mannequin is most succesful. The query now could be which supplier gives probably the most predictable, cost-efficient path to operating AI at manufacturing scale. OpenAI’s cuts are a direct response to that shift, and so they reframe the GPT-5.6 sequence from a premium entry product right into a cost-competitive deployment platform.
Comparative Methods of OpenAI, Google, and Anthropic
Every of the three main frontier suppliers has chosen a unique mechanism to decrease the full price of manufacturing AI. OpenAI is instantly slicing per-token charges. Google, with Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, is pairing decrease costs with reductions in token consumption and gear calls — Gemini 3.6 Flash reportedly makes use of 17% fewer output tokens than Gemini 3.5 Flash on the Synthetic Evaluation Index, with financial savings reaching as much as 65% on some long-horizon engineering duties. Anthropic took a 3rd path: Claude Opus 5 prices the identical $30 mixed per million tokens as Opus 4.8, however delivers near-Fable 5 efficiency — successfully reducing the value per unit of functionality with out transferring the sticker value. Anthropic additionally added an adjustable effort setting permitting builders to commerce reasoning depth for velocity and token financial savings.
All three approaches goal the identical operational metric: the full price of finishing actual manufacturing work, not simply the marketed value of a single token. For enterprise consumers evaluating which platform to standardize on, that distinction now drives the dialog greater than uncooked benchmark scores.
What stays unresolved is whether or not these value cuts are sustainable at scale or symbolize a land-grab second pushed by aggressive strain. Chinese language open-weight fashions, together with choices from DeepSeek and MiniMax, proceed to push the ground decrease on pure token price, whereas closed-model suppliers spend money on larger infrastructure expenditure. Amazon, for its half, raised its 2026 capital expenditure to $220 billion, per CNBC, partly pushed by rising reminiscence prices — a reminder that the economics of manufacturing low-cost AI tokens are nonetheless beneath strain on the infrastructure stage. For now, OpenAI is betting that frontier-quality intelligence at low-cost pricing is a mix the market pays for, even when not everybody on the backside of the value desk can match it.
FAQ
How a lot did OpenAI scale back the value of GPT-5.6 Luna?
OpenAI minimize the value of GPT-5.6 Luna by 80%, reducing the mixed enter and output token price to $1.40 per million tokens — down from a earlier mixed value of $7 per million tokens.
What’s the new pricing technique for OpenAI’s GPT-5.6 Sol mannequin?
OpenAI added a premium Sol Quick mode that gives as much as 2.5 instances throughput at double the price of the Normal mode, charging $70 per million tokens mixed ($10 enter, $60 output). Sol Normal pricing stays unchanged at $35 mixed per million tokens.
How does OpenAI’s Luna pricing examine to Google’s Gemini AI fashions?
Luna’s $1.40 mixed token value is cheaper than Google’s Gemini 3.5 Flash-Lite at $2.80 and considerably beneath Gemini 3.6 Flash at $9 per million tokens mixed.
What workloads are the GPT-5.6 Luna, Terra, and Sol fashions designed for?
Luna targets high-throughput, low-latency duties resembling summarization, classification, and routing. Terra balances functionality and effectivity for common manufacturing workloads. Sol focuses on complicated reasoning, superior coding, multi-step planning, and agentic methods.
Article produced with the help of synthetic intelligence and reviewed by the editorial group.
