Skip to main content

OpenAI's Jalapeño Inference Chip Outperforms Nvidia Blackwell in Early Benchmarks

25 AUGUST 2026·3 MIN READ·15 SOURCES·Official source plus independent coverage

OpenAI has published the first performance results for Jalapeño, its custom artificial intelligence inference chip developed with Broadcom. Operating at lower power than competing systems, the chip delivered significant gains in throughput and latency across multiple open-weight models.

OpenAI's Jalapeño Inference Chip Outperforms Nvidia Blackwell in Early Benchmarks

Key takeaways · 3

  • 01

    Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput.

  • 02

    The 700-watt chip operates at significantly lower power than Nvidia's 1,200 to 1,400-watt systems.

  • 03

    OpenAI plans to deploy the chip in small volumes by late 2026.

Performance and Benchmarks

OpenAI published the first benchmark results for Jalapeño, its custom artificial intelligence inference chip. [1] Tested on the InferenceX benchmark suite, the processor delivered 1.5 to 1.9 times more artificial intelligence work per watt at peak throughput compared to comparison systems. [1] Jalapeño registered both more tokens per user and more throughput per kilowatt than currently available state-of-the-art inference processors. [2]

The chip was tested against three open models: GPT-OSS 120B, DeepSeek R1, and Moonshot AI's one-trillion-parameter Kimi K2.5. [11] Across these models, the processor achieved 1.7 to 3.6 times lower end-to-end latency than Nvidia's GB200 and GB300 systems. [11] At low-latency operating points, the company claims 8.6 to 104.3 times more throughput per kilowatt at the fastest previous time-between-tokens settings for the GB300. [11]

Architecture and Efficiency

OpenAI developed the chip in close collaboration with Broadcom, advancing from initial team hiring to manufacturing tape-out in approximately 16 months. [7] The 700-watt processor achieved its results operating at lower power requirements than Nvidia's GB200 and GB300 systems, which draw between 1,200 and 1,400 watts. [11] The chip features HBM4 memory, positioning it as comparable to flagship graphics processing units from Nvidia and AMD. [7]

The architecture is designed to minimize delays during the prefill and communication phases of inference processing. [2] By keeping model state explicit and local, the system activates specific combinations of compute, memory, and networking for each inference phase to minimize data movement. [2]

Development and Deployment

OpenAI utilized its own artificial intelligence models to accelerate Jalapeño's development process. [1] Earlier generations assisted the team in designing the chip, while the company's latest models are accelerating how the chip is optimized and programmed. [1] The hardware is intended to function as a multigenerational platform, enabling the concurrent development of artificial intelligence products, models, processors, and memory. [2]

The company plans to integrate the hardware into its data centers later this year to support its own artificial intelligence models. [9] Richard Ho, OpenAI's head of hardware, estimated that the processor will deploy in very small volumes at the end of 2026, with more significant deployment following in 2027. [2]

Economic and Operational Impact

Operating at 700 watts allows the company to save money when running data centers, where power represents a key cost. [9] Ho stated that Jalapeño demonstrates performance in both the high-throughput domain, allowing the company to serve many customers more cheaply, and the low-latency domain for faster response times. [9]

For highly interactive workloads, the chip delivered 2.1 to 4.1 times higher performance. [1] For consumers, this combination can translate to faster responses, more responsive Codex sessions, and more reliable access during periods of peak demand. [11] The company stated that developing in-house silicon grants more direct control over model execution and the economics of serving artificial intelligence models. [10]

What it means

By combining high throughput and low latency into a single architecture, OpenAI is attempting to rewrite the traditional hardware tradeoff that forces data centers to optimize for one or the other. This efficiency is critical as power consumption becomes the primary bottleneck for AI scaling. Comparing the 700-watt Jalapeño against Nvidia's 1,200-to-1,400-watt GB200 and GB300 systems highlights a strategic shift toward custom, model-agnostic inference silicon. If the company achieves its 2027 deployment targets, it could significantly alter the margin structure of enterprise AI serving by reducing dependency on expensive, power-hungry generic GPUs. What the sources don't address: how Nvidia's next-generation architectures will compare in power efficiency by the time Jalapeño reaches full-scale deployment in 2027.

The transition to custom inference silicon directly impacts the cost and speed of AI deployment. As models grow, hardware efficiency like Jalapeño's 700W footprint will dictate which platforms can scale profitably.

Why it matters
Story quiz

Test yourself on this story — 3 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 2 September 2026

    Archived

  2. 29 August 2026

    Event evidence refreshed from source cluster.

  3. 27 August 2026

    Event evidence refreshed from source cluster.

  4. 27 August 2026

    Updated with 3 new sources — now corroborated by 14 sources.

  5. 26 August 2026

    Event evidence refreshed from source cluster.

  6. 26 August 2026

    Updated with 6 new sources (industry) — now corroborated by 11 sources.

  7. 25 August 2026

    OpenAI Benchmarks Custom Jalapeño Chip, Beating Nvidia in Early Tests

  8. 25 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.