Skip to main content

OpenAI Benchmarks Custom Jalapeño Chip, Beating Nvidia in Early Tests

25 AUGUST 2026·2 MIN READ·5 SOURCES·Official source plus independent coverage

OpenAI revealed performance benchmarks for its first custom inference chip, Jalapeño, demonstrating significant throughput and latency advantages over current Nvidia processors.

OpenAI Benchmarks Custom Jalapeño Chip, Beating Nvidia in Early Tests

Key takeaways · 3

  • 01

    Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput.

  • 02

    The chip achieved 1.7 to 3.6 times lower end-to-end latency than Nvidia systems.

  • 03

    OpenAI plans to deploy the hardware in very small volumes at the end of 2026.

Hardware Architecture

At the Hot Chips conference, OpenAI detailed Jalapeño, its first custom inference chip created in collaboration with Broadcom over 16 months. [2][4] First announced last October, the hardware is explicitly designed to minimize data movement and communication delays during the prefill phase of processing. [2] By keeping the KV cache local during response generation, Jalapeño activates the specific compute and memory needed for each inference phase. [2]

The processor is part of a broader full-stack strategy where OpenAI co-designs models, serving software, memory, and networking together. [1] Early OpenAI models helped the hardware team design the chip, while current models are accelerating how the company optimizes and programs it. [1]

Performance Metrics

OpenAI evaluated the processor using SemiAnalysis’ InferenceX benchmark. [2] Across models including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput than comparison systems. [1][5] The chip also achieved 1.7 to 3.6 times lower end-to-end latency. [1]

For highly interactive workloads, performance was measured at 2.1 to 4.1 times higher. [1] OpenAI’s head of hardware, Richard Ho, stated the chip serves more AI work per unit of power while returning responses more quickly, as the chip architecture avoids the traditional hardware tradeoff between throughput and latency. [1][2]

Deployment Timeline

Jalapeño, built as a specialized ASIC, produced benchmarks that beat processors from Nvidia, AMD, and Google on multiple top open-weight models. [4] Specifically, the primary benchmark comparisons were made against an Nvidia Blackwell system, where OpenAI claims its new chips can outperform Nvidia processors in tests. [2][3]

OpenAI intends for Jalapeño to serve as the foundation of a multigenerational hardware platform. [1][2] According to Ho, the processor will begin deployment in very small volumes at the end of 2026, with a more significant rollout planned for 2027. [2]

What it means

OpenAI's shift toward custom silicon marks a direct challenge to its reliance on Nvidia Blackwell architectures. By focusing on minimizing communication bottlenecks and co-designing hardware with its own serving software, OpenAI is optimizing specifically for the economics of high-throughput AI inference. These efficiency gains are critical as model serving costs scale across highly interactive workloads. The 16-month development timeline with Broadcom demonstrates how aggressively the AI lab is pursuing hardware independence. What the sources don't address: How much Broadcom is charging for the design collaboration and whether OpenAI will eventually sell or lease chip access to other cloud providers.

Custom inference silicon allows AI developers to optimize power and latency beyond standard GPUs. This efficiency directly impacts the cost and speed of delivering AI applications to end users.

Why it matters
Daily session

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 25 August 2026

    OpenAI Benchmarks Custom Jalapeño Chip, Beating Nvidia in Early Tests

  2. 25 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.