OpenAI slashed the price of its cheapest and fastest model, GPT-5.6 Luna, by 80% on July 30. That’s about three weeks after the GPT-5.6 line went to users.
Chinese open-weight models continue to pull token traffic from US labs on price, especially for developers running big batches of AI work.
OpenAI announced new pricing for Luna at $0.20 per million input tokens and $1.20 per million output tokens. Prices for Terra, the middle tier aimed at everyday tasks, slid 20% to $2 and $12 per million tokens. Sol, the best coding model from OpenAI, remained at $5 per million input tokens and $30 per million output tokens.
The company has swapped out its previous Priority Processing tier for a new “Fast Mode” that it claims can run Sol up to 2.5 times faster than standard processing, while maintaining the same level of model intelligence.
Fast mode costs twice as much, so Sol under the fast tier costs $10 per million input tokens and $60 per million output tokens. The change is backward compatible, and developers who already used the priority tier will be automatically migrated and won’t need to change their integrations, according to OpenAI.
The new pricing also changes how paid plans meter usage. Terra and Luna will be burning fewer credits in ChatGPT Work and Codex, but the prices of subscriptions and quota budgets remain unchanged.
OpenAI says some of the savings were due to work done by Sol. The model rewrote production kernels, ran experiments on token generation, and monitored training runs, intervening when something went wrong, the company said, all under human supervision. Sol’s experiments increased the efficiency of token generation by more than 15%, and that kernel work lowered the cost of serving the model by 20%.
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are… pic.twitter.com/rFhK7XKedp
— OpenAI (@OpenAI) July 30, 2026
The company claims Luna is comparable to frontier-class models from a year ago, at about six cents on the dollar per task and about nine times the speed. Additionally, OpenAI said Luna bested Anthropic’s Fable 5 on its internal Agents’ Last Exam benchmark at an estimated cost per task about 99% lower, a comparison the company has not released underlying data for.
Hoda Noorian of Notion says the firm’s tests found that Terra was providing quality that was on par with that of GPT-5.5, at half the cost per task, and 60% less time.
Cost is becoming the thing that counts for many developers. As reported by Cryptopolitan, Chinese open-weight models beat US systems on OpenRouter in February 2026, processing more than three times the American token volume. There are labs like DeepSeek, Zhipu, and Moonshot that offer downloadable weights instead of charging per token.
In the comparison, Zhipu’s GLM 5.2 costs $1.40 and $4.40 per million tokens, versus $5 and $25 for Anthropic’s Opus 4.8. Kimi K3 from Moonshot is also less expensive than GPT-5.6 Sol and Fable 5.
Uber exhausted its 2026 AI coding budget in April, and 29% of senior leaders were unable to monitor or manage the costs of operating their AI systems, according to a survey of 2,145 leaders by KPMG.
Earlier this month, developers, including Matt Shumer of OthersideAI, said Sol had deleted files without being asked to, while Bruno Lemos said the model had wiped his production database.
Two weeks before the Sol release, OpenAI warned that the model might be “overly agentic” and interpret instructions too permissively. Anyone pointing Sol at real systems should narrow its permissions, keep backups, and roll out in stages, OpenAI said.
If you’re reading this, you’re already ahead. Stay there with our newsletter.