Michele Catasta, Replit’s president and head of AI, supplied the testimonial OpenAI chose to publish with its new GPT-5.6 prices. Catasta wrote, “GPT-5.6 Luna is the closest we’ve come to intelligence too cheap to meter. I’ve never seen a model this affordable be this powerful, it’s unlocking use cases for Replit we didn’t expect to build for a long time.”
Luna’s API price fell 80 percent.
OpenAI made the cut on July 30, three weeks after launching GPT-5.6, Axios reported. The report noted that price cuts usually arrive months after a launch rather than three weeks and that cheaper Chinese open-weight models have pressed OpenAI and Anthropic to justify their higher costs.
What Changed
- OpenAI cut GPT-5.6 Luna's API price 80 percent on July 30, to $0.20 per million input tokens and $1.20 per million output.
- Terra fell 20 percent to $2 and $12 per million tokens. Sol's standard price did not change.
- OpenAI credits efficiency work, including Sol rewriting production kernels for a 20 percent serving-cost cut, all self-reported.
- OpenAI's adjusted gross margin was about 33 percent in the first quarter of 2026, against a 46 percent internal target.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
What the cut actually costs
As of July 30, 2026, Luna costs $0.20 per million input tokens and $1.20 per million output tokens. The previous rates were $1 and $6. OpenAI also lowered Terra to $2 per million input tokens and $12 per million output tokens, from $2.50 and $15. Sol’s standard price did not change.
At the rates available July 30, Anthropic’s Claude Sonnet 4.6 cost $3 per million input tokens and $15 per million output tokens.
OpenAI’s July 29 post presented Luna as inexpensive at several reasoning settings on the Artificial Analysis Intelligence Index v4.1. At its maximum setting, the chart placed Luna at a score of 51.24 and an estimated $0.0571 per task. GLM-5.2 Max scored 51.09 at $0.2692 per task in the same presentation. The chart appeared inside OpenAI’s own post. The material released with the price change contains no comparison OpenAI did not itself present.
The July 30 ZeroHedge account cited Artificial Analysis’s cost-per-task index, where Moonshot AI’s Kimi K3 completed tasks for 94 cents against $1.04 for OpenAI’s Sol. Kimi K3 is a 2.8 trillion parameter open-weight model released this month, which customers can run on their own infrastructure, and it has beaten proprietary systems on benchmarks. DeepSeek V4 Pro ran at four cents per task, according to the account.
Where OpenAI says efficiency came from
In its July 29 post, OpenAI attributed the discounts to work across its models and serving systems. The company said Sol autonomously rewrote production kernels in a human-led process, reducing end-to-end serving costs by 20 percent. It also reported that experiments designed and run by Sol increased token-generation efficiency by more than 15 percent.
Those figures are OpenAI’s own accounting of its own systems. The company released no independent measurement of them alongside the price change. The same holds for OpenAI’s claim that Luna beats Fable 5 on Agents’ Last Exam at an estimated cost per task nearly 99 percent lower.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Fast mode replaced Priority Processing for Sol on July 30, running up to 2.5 times faster than Standard at twice the price.
The customer bill behind price pressure
Sam Altman had previously called costs “a huge issue,” Reuters reported. His comment came as enterprise buyers grew reluctant to approve large AI budgets without clearer returns and long-running agents consumed more tokens.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Telnyx had been spending roughly $100,000 a day on Anthropic, according to the documented case in the July 28 account by Forbes senior contributor Peter Cohan. After switching to Chinese startup Z.AI, the reported cost was about $100 per agent per day. The account as captured does not state how many agents Telnyx ran or whether the workloads and results were directly comparable.
Falling token prices squeeze expensive model providers, Cohan argues. His case identifies chipmakers and Chinese labs as the beneficiaries. Token costs fell from about $60 per million tokens in 2021 to roughly $0.06 by July 2026, according to figures he cites. By mid-2026, DeepSeek, Qwen and GLM had captured 46 percent market share, according to another figure in his account.
What the cuts leave unanswered
OpenAI entered this price cycle with an adjusted gross margin of about 33 percent in the first quarter of 2026, according to Control Plane’s account. The internal target was 46 percent. Revenue reached $5.7 billion in the quarter, and the operating loss was $3.7 billion. Inference costs were cited as the reason for the margin decline. The account said those costs quadrupled in 2025. OpenAI budgeted $32 billion for training during 2026.
The company has not said what the Luna and Terra reductions will do to that margin. Control Plane’s account put OpenAI’s first-quarter 2026 cash reserves at $73 billion and projected its full-year inference costs at roughly $14 billion. Control Plane placed those figures in the context of a potential public listing.
OpenAI filed a confidential draft S-1 with the Securities and Exchange Commission on June 8, 2026, after considering a September listing at a valuation of as much as $1 trillion. Advisers later counseled a delay to 2027, Reuters reported in late June.
Frequently Asked Questions
How much does GPT-5.6 Luna cost now?
As of July 30, 2026, Luna costs $0.20 per million input tokens and $1.20 per million output tokens. The previous rates were $1 and $6, an 80 percent reduction that took effect three weeks after the model launched.
Did every GPT-5.6 model get cheaper?
No. Terra dropped 20 percent, to $2 per million input tokens and $12 per million output, from $2.50 and $15. Sol's standard price was unchanged. Fast mode replaced Priority Processing for Sol on July 30, running up to 2.5 times faster than Standard at twice the price.
Why did OpenAI cut prices three weeks after launch?
OpenAI attributes the cuts to efficiency work across its models and serving systems. Axios reported that price cuts usually arrive months after a launch rather than three weeks, and that cheaper Chinese open-weight models have pressed OpenAI and Anthropic to justify their higher costs.
How does Luna compare with cheaper rivals?
On the Artificial Analysis cost-per-task index cited in a July 30 ZeroHedge account, Moonshot AI's Kimi K3 completed tasks for 94 cents against $1.04 for OpenAI's Sol. DeepSeek V4 Pro ran at four cents per task. Kimi K3 is an open-weight model customers can run on their own infrastructure.
What do the cuts do to OpenAI's margin?
The company has not said. Control Plane put OpenAI's first-quarter 2026 adjusted gross margin at about 33 percent against a 46 percent internal target, with revenue of $5.7 billion and an operating loss of $3.7 billion, and attributed the decline to inference costs.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR