OpenAI details GPT-5.6 efficiency improvements

GPT-5.6 models cut AI costs while boosting speed and coding performance See how Sol, Terra, and Luna trim latency and price across OpenAI's stack

OpenAI says its GPT5.6 model family is designed to balance capability and cost across a range of tasks. The company says GPT5.6 Sol, its flagship reasoning model, performs better than Claude Fable 5 on the Artificial Analysis Coding Agent Index while costing less than half as much, while Terra matches GPT5.5 intelligence benchmarks at about half the price and Luna is positioned as the fastest, least expensive option. In a July 29 engineering post, OpenAI said many of the gains came from changes across its stack, including model training, inference, and its agentic harness used in Codex and ChatGPT Work. The company highlighted work on routing, scheduling, kernel optimization, caching, speculative decoding, and context management as ways to reduce compute use and improve throughput. OpenAI also said GPT5.6 Sol helped automate parts of the optimization process, including analyzing production traffic, rewriting kernels, and tuning speculator models. The company said these changes reduced serving costs and improved tokengeneration efficiency, and it expects further efficiency gains to continue lowering costs and latency for users and customers.