OpenAI's custom inference chip, Jalapeño, has posted benchmark results the company claims are industry-leading in both speed and power efficiency for AI inference workloads.

The chip is purpose-built for modern large language models, targeting two metrics that matter most in production deployment: throughput, meaning how many requests it can handle simultaneously, and latency, meaning how fast individual responses complete. Power efficiency gains on custom silicon typically translate directly to lower operating costs at scale, which is why Google's TPUs and Amazon's Trainium exist. OpenAI is now in that game.

The full post is worth reading for the specific benchmark numbers, the architectural decisions that drove the efficiency gains, and what Jalapeño's existence signals about OpenAI's long-term dependence on Nvidia.

[READ ORIGINAL →]