Advancing the Price-Performance Frontier with GPT-5.6
OpenAI is cutting API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, effective July 30, with Luna priced at $0.20 per million input tokens and $1.20 per million output tokens, and Terra at $2 per million input tokens and $12 pe…
OpenAI is cutting API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, effective July 30, with Luna priced at $0.20 per million input tokens and $1.20 per million output tokens, and Terra at $2 per million input tokens and $12 per million output tokens. The price reductions stem from model-level and infrastructure efficiency gains, including a 20% reduction in end-to-end serving costs and a 15% improvement in token-generation efficiency, partially achieved through GPT-5.6 Sol autonomously optimizing production kernels and running experiments. OpenAI is also replacing its Priority Processing tier with Fast mode for GPT-5.6 Sol, delivering up to 2.5× faster speeds at twice the standard price with no change in intelligence. The updates are designed to make high-volume AI workloads economically viable at scale while preserving access to frontier-speed processing for time-sensitive enterprise tasks.
- 01OpenAI is cutting API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, effective July 30, with Luna priced at $0.20 per million input tokens and $1.20 per million output tokens, and Terra at $2 per million input tokens and $12 per million output tokens.
- 02The price reductions stem from model-level and infrastructure efficiency gains, including a 20% reduction in end-to-end serving costs and a 15% improvement in token-generation efficiency, partially achieved through GPT-5.6 Sol autonomously optimizing production kernels and running experiments.
- 03OpenAI is also replacing its Priority Processing tier with Fast mode for GPT-5.6 Sol, delivering up to 2.5× faster speeds at twice the standard price with no change in intelligence.
- 04The updates are designed to make high-volume AI workloads economically viable at scale while preserving access to frontier-speed processing for time-sensitive enterprise tasks.