Michele Catasta described Replit’s plans by naming work that had been left for later. Replit, which builds software for writing and running code, had use cases it did not expect to build for a long time, he wrote in a testimonial published by OpenAI. Catasta, the company’s president and head of AI, called GPT-5.6 Luna “the closest we've come to intelligence too cheap to meter.”

On July 30, OpenAI cut Luna’s launch price by 80 percent.

The reduction arrived three weeks after GPT-5.6 launched and covered Luna and the mid-range Terra model, while Sol stayed at its existing rate, Axios reported. The change put a much smaller charge on the routine model calls that accumulate inside coding agents and other business software.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Luna pricing and tokens

As of July 30, 2026, Luna costs $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6. A token is the billing unit OpenAI counts when a model reads a prompt or produces a reply. It may be a word or a piece of one, so the price sheet does not describe a finished report, a solved coding problem or an answered customer request. If a system reads one million tokens and writes one million back, the model charge is $1.40 at Luna’s new rates, before any other computing or software costs.

An inference bill is the running tab for using a trained model. Each prompt adds input tokens; each response adds output tokens, and longer agent jobs can repeat that exchange many times while calling tools or checking their own work. An 80 percent cut therefore reduces the charge on every Luna request, though the total saving depends on how much text the application sends and receives.

Terra’s July 30 rate moved to $2 per million input tokens and $12 per million output tokens, from $2.50 and $15. Sol received no price reduction. OpenAI instead introduced Fast mode for Sol, which it says runs up to 2.5 times faster than Standard processing at twice the price.

Replit’s delayed work

Catasta sits on the customer side of that arithmetic. His reference to delayed Replit use cases gave OpenAI a concrete beneficiary for the cut, though the testimonial did not identify the projects or publish their expected token use. “I've never seen a model this affordable be this powerful, it's unlocking use cases for Replit we didn't expect to build for a long time,” he wrote.

In its July 29 post, OpenAI credited changes in models, serving systems and the software that connects agents to tools and context. The company described a human-led process in which Sol rewrote and optimized production kernels, then designed and ran hundreds of experiments. OpenAI attributed a 20 percent reduction in end-to-end serving cost and a gain of more than 15 percent in token-generation efficiency to that work. It also claimed Luna beat Fable 5 on Agents’ Last Exam at a cost per task nearly 99 percent lower.

Those figures are OpenAI’s own measurement of its own systems. The company released no independent measurement of the 20 percent serving-cost reduction, the 15 percent token-generation gain or the nearly 99-percent-lower comparison.

Peter Cohan’s cost case

Peter Cohan approached falling token prices from the supplier’s side. In a July 28 Forbes column, the senior contributor traced the price of one million tokens from $60 in 2021 to roughly $0.06 by July 2026 and argued that expensive model providers face thinner profits as buyers route routine work to cheaper systems.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

His customer example was Telnyx. Telnyx had been spending roughly $100,000 a day with Anthropic, Cohan reported on July 28, before shifting work to Chinese startup Z.AI at about $100 per agent per day. The figures use different bases, and the source gives no agent count for calculating Telnyx’s total replacement bill. Cohan presented the switch as evidence that customers now examine the cost of each task.

Sam Altman had already called costs “a huge issue,” Reuters reported. The OpenAI chief executive said the company would help customers get more value, addressing the budget approvals that enterprise customers had grown reluctant to make without clear returns.

OpenAI’s margin target

OpenAI entered the price cut with an adjusted gross margin of about 33 percent in the first quarter of 2026, against an internal target of 46 percent, Control Plane reported. The publication attributed the decline to higher inference costs and reported that those costs quadrupled during 2025. Its estimate for 2026 was approximately $14 billion, alongside a $32 billion training budget.

No source says that margin pressure caused the Luna and Terra reductions. OpenAI presented the lower rates as the result of engineering gains, and the company’s adjusted gross margin covers more than the price of one model family.

OpenAI submitted a confidential draft S-1 to the Securities and Exchange Commission on June 8, 2026, while working toward a possible September 2026 listing at a valuation of as much as $1 trillion. Reuters reported in late June 2026 that advisers had counseled delaying the offering until 2027.

Frequently Asked Questions

How much does GPT-5.6 Luna cost now?

As of July 30, 2026, Luna costs $0.20 per million input tokens and $1.20 per million output tokens. The previous rates were $1 and $6, an 80 percent reduction that took effect three weeks after the model launched.

Did every GPT-5.6 model get cheaper?

No. Terra dropped 20 percent, to $2 per million input tokens and $12 per million output, from $2.50 and $15. Sol's standard price was unchanged. Fast mode replaced Priority Processing for Sol on July 30, running up to 2.5 times faster than Standard at twice the price.

Why did OpenAI cut prices three weeks after launch?

OpenAI attributes the cuts to efficiency work across its models and serving systems. Axios reported that price cuts usually arrive months after a launch rather than three weeks, and that cheaper Chinese open-weight models have pressed OpenAI and Anthropic to justify their higher costs.

How does Luna compare with cheaper rivals?

On the Artificial Analysis cost-per-task index cited in a July 30 ZeroHedge account, Moonshot AI's Kimi K3 completed tasks for 94 cents against $1.04 for OpenAI's Sol. DeepSeek V4 Pro ran at four cents per task. Kimi K3 is an open-weight model customers can run on their own infrastructure.

What do the cuts do to OpenAI's margin?

The company has not said. Control Plane put OpenAI's first-quarter 2026 adjusted gross margin at about 33 percent against a 46 percent internal target, with revenue of $5.7 billion and an operating loss of $3.7 billion, and attributed the decline to inference costs.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Moonshot Releases Kimi K3 Weights as Amodei Rejects Open-Weight Ban
Moonshot AI released the weights for its Kimi K3 model on Monday, a system with 2.8 trillion parameters. Developers can now download, modify and self-host K3, which the report described as the world's
Anthropic's Opus 5 Cut Tokens 17% but Cost More to Benchmark Than Opus 4.8
Anthropic released Claude Opus 5 on July 24 with per-token API rates unchanged from Opus 4.8, the company announced. Artificial Analysis reported that its independent Intelligence Index run cost $3,83
Palo Alto Networks CEO Says AI Token Costs Must Fall Up to 90%
Palo Alto Networks CEO Nikesh Arora said on CNBC on Thursday that AI token costs need to fall as much as 90% to support large-scale enterprise adoption. He called OpenAI CEO Sam Altman’s claim that th
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai