Jalapeño's first results show industry-leading speed and efficiency in AI inference
OpenAI has published benchmark results for Jalapeño, its first custom AI inference chip, showing industry-leading performance across throughput, power efficiency, and latency. Tested against publicly available models including GPT-OSS 12…
- 01Tested against publicly available models including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T using the InferenceX benchmark from SemiAnalysis, Jalapeño delivered 1.5–1.9x more AI work per watt at peak throughput and 1.7–3.6x lower end-to-end latency than competing commercially available systems.
- 02For highly interactive workloads, performance gains reached 2.1–4.1x.
- 03The chip was developed in nine months using AI-assisted design, with AI-generated kernel implementations running 1.5–1.8x faster than human-expert-written code on selected model blocks.
- 04OpenAI plans to deploy Jalapeño within its own infrastructure by end of year, with second- and third-generation designs already in development, positioning the chip as a tool to improve operating leverage by growing useful AI output faster than cost-to-serve.
OpenAI has published benchmark results for Jalapeño, its first custom AI inference chip, showing industry-leading performance across throughput, power efficiency, and latency. Tested against publicly available models including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T using the InferenceX benchmark from SemiAnalysis, Jalapeño delivered 1.5–1.9x more AI work per watt at peak throughput and 1.7–3.6x lower end-to-end latency than competing commercially available systems. For highly interactive workloads, performance gains reached 2.1–4.1x. The chip was developed in nine months using AI-assisted design, with AI-generated kernel implementations running 1.5–1.8x faster than human-expert-written code on selected model blocks. OpenAI plans to deploy Jalapeño within its own infrastructure by end of year, with second- and third-generation designs already in development, positioning the chip as a tool to improve operating leverage by growing useful AI output faster than cost-to-serve.
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.