Thu 06 Aug 2026 / 09:46 ET
Kernel
Hardware 3 min read

OpenAI GPT-5.6 price cuts sharpen the fight for everyday AI work

OpenAI cut Luna by 80% and Terra by 20% while charging extra for faster Sol processing, giving enterprise buyers more reason to split work across models.

Felix Aranda

By Felix Aranda / Silicon Editor

OpenAI GPT-5.6 price cuts sharpen the fight for everyday AI work
img: Tom's Hardware

OpenAI GPT-5.6 price cuts have made its cheaper API models markedly less expensive: on July 30, the company reduced the price of Luna by 80% and Terra by 20%. At the same time, OpenAI introduced a faster processing tier for its flagship Sol model that costs twice the standard rate. The combination is a concrete sign of tighter competition for routine, high-volume AI work, even if it does not prove a market-wide collapse in provider margins.

OpenAI calls Luna its fastest and most affordable model, and Terra its balanced option for everyday work. The company said the reduced pricing also changes how Luna and Terra usage is counted under paid Codex and ChatGPT Work subscriptions.

For Sol, OpenAI replaced Priority Processing with Fast mode. It says the new option can provide up to 2.5 times the speed of standard processing without changing the model's intelligence, for double the price. That leaves customers with an explicit latency bill: pay less where response time is less urgent, or pay more for faster output.

Why are enterprises using multiple AI models?

Companies are increasingly matching a model to a job rather than sending every request to the priciest available system. The Wall Street Journal reported in July that businesses had begun adding lower-priced models, including Chinese offerings, alongside products from OpenAI and Anthropic. The premise is mundane and sensible: routine tasks may not require a flagship model.

That approach is called model routing. A company can reserve a stronger model for work where errors carry more consequence, use a lower-cost model for repetitive steps, and choose a faster tier when delay matters. A large language model produces text in tokens, but the listed price per token is only one part of a deployment decision. The number of tokens, tool calls, retries and model passes required to complete a workflow can differ by model and task.

OpenAI argues its cuts reflect engineering gains rather than price alone. It said changes to model design, inference systems, routing, production software and context management reduce the time, tokens and cost needed to produce results. It also said internal kernel work cut end-to-end serving costs by 20% and experiments improved token-generation efficiency by more than 15%. Those figures are OpenAI's claims, not independently audited results.

The company says customers should weigh stakes, error costs, urgency, volume and required quality for each stage of a workflow. That advice also describes the pressure facing providers. As buyers gain more credible alternatives for ordinary tasks, model vendors have to compete on the combined result of capability, speed and price.

For now, the evidence establishes a developing pricing squeeze rather than a settled verdict on profitability. Lower API list prices do not, by themselves, show what providers spend to serve a request, what customers spend overall, or whether a particular model earns money. OpenAI's decision to discount Luna and Terra while charging more for fast Sol access shows the market is fragmenting by workload instead of converging on one universal AI price.

This story draws on original reporting from Tom's Hardware.

More Hardware/

view all ↗