The Information Machine

Advancing the price-performance frontier with GPT-5.6

OpenAI Blog · 2026-07-30

OpenAI cuts GPT-5.6 Luna prices by 80% and Terra prices by 20%, enabled by GPT-5.6 Sol autonomously optimizing its own production GPU kernels and token-generation experiments to reduce serving costs by 20% and increase efficiency by 15%.

Open original ↗

Appears in

Extraction

Topics: ai-pricinginference-efficiencygpt-5-6ai-infrastructurerecursive-self-improvement

Claims

  • GPT-5.6 Luna prices dropped 80% and GPT-5.6 Terra prices dropped 20% starting July 30, 2026, with Luna costing $0.20 per million input tokens.
  • GPT-5.6 Sol autonomously rewrote production GPU kernels within a human-led process, cutting end-to-end serving costs by 20% and increasing token generation efficiency by more than 15%.
  • Luna delivers performance comparable to frontier-class models from one year ago at approximately 6 cents on the dollar per task and nearly nine times the speed.
  • Fast mode for GPT-5.6 Sol replaces Priority Processing in the API, delivering up to 2.5x faster speeds at twice the standard price with no change in model intelligence.
  • The self-optimization feedback loop is expected to accelerate future efficiency gains as increasingly capable models take on more of the infrastructure improvement work.

Key quotes

GPT-5.6 Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training, intervening when problems arose.
The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%.
Luna delivers performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task, and at nearly nine times the speed.