Building abundant intelligence
OpenAI Blog · 2026-07-31
OpenAI announces an 80% price cut to GPT-5.6 Luna and describes a full-stack infrastructure strategy in which AI-assisted engineering, smarter context management, and demand-driven investment form a self-reinforcing cycle intended to make powerful intelligence progressively cheaper and more accessible.
Appears in
Extraction
Topics: ai-pricingai-infrastructuremodel-efficiencyagentic-aiopenai-strategy
Claims
- OpenAI reduced GPT-5.6 Luna input pricing by 80% to $0.20 per million tokens and Terra by 20%, while launching Sol Fast at 2.5x speed for twice the price.
- GPT-5.6 Sol helped optimize OpenAI's own production serving software, reducing end-to-end serving costs by 20% and improving speculative decoding token-generation efficiency by more than 15%.
- Improvements to context management and retained reasoning raised GPT-5.6 Sol's ARC-AGI-3 score from 13.3% to 38.3% using six times fewer output tokens, without modifying the model itself.
- OpenAI's models now serve over one billion active users and two million businesses, with users sending roughly 50% more messages per day after six months.
- Agentic work through Codex accounts for 99.8% of OpenAI's weekly output tokens, marking a fundamental shift from question-answering to task completion.
Key quotes
Better intelligence drives broader adoption. Broader adoption supports more investment. More investment improves intelligence and efficiency. That is the cycle we are building.
The model did not change. The surrounding system did.
Our goal is not simply more compute, bigger models, or lower token prices. It is more useful intelligence within reach.