The Performance Metrics Of OpenAI’s Jalapeño Chip In AI Benchmarks
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Performance Metrics Of OpenAI’s Jalapeño Chip In AI Benchmarks on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

OpenAI has released initial performance metrics for its Jalapeño inference chip, showing notable efficiency and latency improvements against NVIDIA’s Blackwell GPUs in AI benchmarks. The results are based on OpenAI’s own testing and are not yet independently verified.

OpenAI has released the first publicly measured performance results for its Jalapeño inference chip, indicating substantial efficiency and latency improvements over NVIDIA’s Blackwell systems in AI inference benchmarks. These results are significant as they highlight OpenAI’s progress in custom hardware development for AI workloads, though they are based on vendor-provided measurements and have not yet been independently verified.

According to OpenAI, Jalapeño achieved between 1.5 to 1.9 times higher performance per watt and 1.7 to 3.6 times lower latency across three benchmarked models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. These tests were conducted on OpenAI’s own hardware setup, comparing Jalapeño against NVIDIA’s Blackwell-based systems, specifically the GB200 and GB300 models.

OpenAI emphasized that Jalapeño is a dedicated inference ASIC optimized for specific AI workloads, contrasting with NVIDIA’s general-purpose GPUs. The performance metrics were measured using InferenceX, a public benchmark that evaluates the entire inference process, including prompt prefill and token generation. The results suggest that Jalapeño can deliver more efficient AI inference, especially in latency-sensitive applications, which are critical for AI agents and interactive systems.

However, these measurements are vendor-reported, and Jalapeño has not yet been deployed in OpenAI’s production environment. The chip is still undergoing qualification, with deployment anticipated by the end of 2024. The results are promising but require independent validation to confirm the claims made by OpenAI.

At a glance
reportWhen: announced March 2024
The developmentOpenAI published first measured results for its Jalapeño inference chip, revealing promising performance metrics against NVIDIA systems in AI benchmarks.
AI DISPATCH · REALITY CHECKOpenAI Jalapeño · part 1 of 2 · 25 Aug 2026
The numbers are strong — and they’re the vendor’s
Jalapeño’s First Results: Read the Metric, Not the Headline

OpenAI’s first custom inference chip posts real per-watt wins on a public benchmark — measured by OpenAI, on the metric OpenAI chose, against NVIDIA only, on a chip not yet deployed.

1.5–1.9×
More AI work per watt (peak)
1.7–3.6×
Lower end-to-end latency
2.1–4.1×
Higher on interactive workloads
Per-watt inference — InferenceX (SemiAnalysis), OpenAI-run
Three external models, all vs NVIDIA Blackwell

Normalized by published TDP: Jalapeño 700W (measured ≤550W) vs GB200 1,200W / GB300 1,400W. Peak throughput per kW — higher is better.

GPT-OSS 120B mixed TPS / kW
vs GB200 · ~1.9×
Jalapeño
85.4k
GB200
45.0k
DeepSeek R1 670B mixed TPS / kW
vs GB300 · ~1.7×
Jalapeño
19.6k
GB300
11.8k
Kimi K2.5 1T mixed TPS / kW · largest tested
vs GB300 · ~1.5×
Jalapeño
18.2k
GB300
11.9k
Read the metric — three things the headline hides
~“Per watt” is a choice. Defensible for datacenter economics, but it structurally favors the lower-power part. Per-chip or per-dollar would read differently.
!ASIC vs general-purpose GPU. Blackwell trains and infers; Jalapeño does one job. Beating a GPU on inference-per-watt is why you build an ASIC — not a full verdict on the GPU.
iVendor-reported, not yet deployed. OpenAI’s own measurements; ships inside OpenAI by year-end, qualification ongoing. Ignore the 50–100× “at previous TBT” cherry — it’s one narrow operating point.

Implications of Jalapeño’s Performance Gains

The reported improvements in efficiency and latency could have substantial impacts on AI infrastructure, reducing operational costs for large-scale AI deployment and enabling faster, more responsive AI agents. Custom inference chips like Jalapeño are designed to optimize specific workloads, potentially reshaping hardware choices in data centers focused on AI services. If independently verified, these results could accelerate adoption of specialized hardware, impacting competitors and the broader AI ecosystem.

Furthermore, the architectural approach—focusing on minimizing data movement and optimizing for both prompt prefill and token decode phases—demonstrates a shift toward workload-specific hardware design. This could influence future hardware development strategies across the industry, emphasizing balanced, adaptable accelerators for AI inference.

Distributed AI Systems: A practical guide to building scalable training, inference, and serving systems for production AI

Distributed AI Systems: A practical guide to building scalable training, inference, and serving systems for production AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Hardware and Benchmarking

OpenAI has historically relied on NVIDIA GPUs for training and inference, with recent efforts to develop custom hardware gaining attention. The company announced Jalapeño earlier this year as part of its strategy to improve inference efficiency. Prior benchmarks have largely focused on GPU performance, with less emphasis on specialized inference ASICs. The release of these initial measurements marks a significant step in demonstrating the potential of custom hardware for AI workloads.

OpenAI's testing was conducted using publicly available benchmarks like InferenceX, which measures the entire request cycle for AI models. These benchmarks are increasingly used to compare hardware performance in real-world AI serving scenarios, providing more practical insights than raw throughput or FLOPS alone. The results, while promising, are still preliminary and based on internal testing conditions.

Amazon

AI inference chips

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Nature of Performance Data

The performance results are based on OpenAI's own measurements and have not been independently verified by third parties. Jalapeño has yet to be deployed in a production environment, and the testing conditions may differ from real-world usage. The full capabilities and performance in diverse workloads remain to be seen.

Additionally, the comparison is limited to NVIDIA's Blackwell systems, and no data is available yet on how Jalapeño performs relative to other hardware providers like AMD or Google. The long-term stability and scalability of Jalapeño are also still under assessment.

Amazon

AI hardware accelerator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Validation and Deployment

OpenAI plans to continue testing Jalapeño in real-world data center environments and aims to begin deploying the chip within its infrastructure by the end of 2024. Independent benchmarking agencies are expected to evaluate the chip's performance, which will be critical for confirming OpenAI's claims.

Further technical disclosures and detailed performance data are anticipated as the chip moves toward production, providing a clearer picture of its capabilities and potential industry impact. OpenAI's ongoing development efforts will also focus on optimizing the architecture for broader AI workloads and scaling performance.

Amazon

custom AI inference ASIC

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does Jalapeño compare to NVIDIA GPUs in AI inference?

Based on OpenAI's internal measurements, Jalapeño shows between 1.5 to 1.9 times higher performance per watt and 1.7 to 3.6 times lower latency than NVIDIA's Blackwell systems in tested benchmarks. However, these results are preliminary and vendor-reported.

Is Jalapeño already deployed in OpenAI's infrastructure?

No, Jalapeño is still undergoing qualification and has not yet been deployed. Deployment is expected by the end of 2024, pending further testing and validation.

What makes Jalapeño different from general-purpose GPUs?

Jalapeño is a dedicated inference ASIC designed specifically for AI workloads, focusing on minimizing data movement and optimizing for both prompt prefill and token decode phases. This specialization aims to improve efficiency and latency compared to GPU-based inference.

Will independent benchmarks confirm these performance claims?

That remains to be seen. The current results are from OpenAI's own testing, and independent evaluations will be necessary to verify the performance improvements claimed.

What is the significance of performance per watt as a metric?

Performance per watt measures efficiency, which is especially important for data centers managing power costs. OpenAI emphasizes this metric because it reflects real-world operational savings, though it may favor lower-power hardware in comparisons.

Source: ThorstenMeyerAI.com

You May Also Like

AI’s Impact On The China Open-Weight Window: A New Global Arena

Analysis of recent US and Chinese policy moves shaping the global AI open-weight landscape and implications for innovation and security.

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

GPT-5.6 Sol Ultra has produced a formal proof of the Cycle Double Cover Conjecture, marking a significant breakthrough in graph theory research.

Mistral Forge: Owning the Model, Not Just Renting the API

Mistral’s Forge enables organizations to build and operate their own AI models, moving beyond API rentals to full ownership and control, with specific use cases in mind.

Disk Is the Contract: Inside Threlmark’s Local-First Architecture

Discover how Threlmark’s local-first system uses disk as the ultimate source of truth, enabling offline resilience, privacy, and seamless multi-device workflows.