
OpenAI Jalapeño Benchmarks: 1.9x Work per Kilowatt
OpenAI published the first Jalapeño inference benchmarks: a 700W part claiming up to 1.9x throughput per kilowatt and 3.6x lower end-to-end latency.
When OpenAI and Broadcom unveiled the Jalapeño inference chip earlier this summer, the interesting part was the development timeline — a reticle-sized ASIC designed from scratch in nine months. What was missing was numbers. On August 25, 2026, OpenAI published the first benchmark results, and the headline claim is that a single architecture delivered both higher throughput and lower latency at the same time, which is normally a trade-off you have to pick a side of.
- OpenAI reports 1.5x to 1.9x more work per watt at peak throughput versus its comparison systems
- End-to-end latency came in 1.7x to 3.6x lower, with 2.1x to 4.1x higher performance on the most interactive workloads
- Jalapeño is rated at 700W but measured at or below 550W sustained on the workloads tested
- Testing used InferenceX, the public serving benchmark from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T
A note on attribution: these are OpenAI's own published figures for its own silicon, run on a third-party benchmark. They are worth taking seriously and worth reading as vendor-reported until independent labs publish their own runs.
What Did the Jalapeño Benchmarks Actually Measure?
OpenAI ran the chip on InferenceX, a public benchmark from SemiAnalysis that measures the whole path of serving an AI request rather than a single kernel in isolation. That choice matters. Peak matrix throughput is easy to quote and hard to feel; end-to-end serving latency is what a user actually experiences when a model starts streaming a reply.
Three models were used as the workload set: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. That spread covers a mid-size open-weight model, a reasoning model, and a trillion-parameter mixture-of-experts system, which is a reasonable sample of what LLM inference looks like in production in 2026.
How Fast Is It Compared With Nvidia Blackwell?
Against the comparison systems — Nvidia's GB300 Blackwell configuration in the published charts — OpenAI reports:
- 1.5x to 1.9x more work per watt at peak throughput across all three models
- 1.7x to 3.6x lower end-to-end latency
- 2.1x to 4.1x higher performance on the most interactive, low-concurrency workloads
- On Kimi K2.5 1T, 1.56 seconds of latency versus 5.31 seconds for the GB300
- On DeepSeek R1, more than 700 tokens per second per user at a concurrency of one
The power comparison is the cleanest part of the story: a 700W Jalapeño part measured against a 1,400W flagship GPU, delivering up to 1.9x the throughput per kilowatt. And the chip did not run at its rating — sustained draw stayed at or below 550W on the tested workloads, which is the kind of headroom that shows up later as denser racks rather than as a bigger number on a slide.
Why Does Performance per Watt Decide AI Infrastructure?
Because power, not silicon, is the binding constraint on modern inference build-outs. A data centre hall has a fixed electrical envelope, and once you have filled it, the only way to serve more requests is to get more work out of each watt. This is the same lens we applied to Nvidia's Vera Rubin NVL72 tokens-per-watt figures and to Meta's MTIA 400 FP4 accelerator, and it is why every serious operator is now designing to a perf-per-watt target rather than a peak-FLOPS one.
There is a second reason the efficiency number is the one to watch: an inference-only accelerator gives up generality in exchange for exactly this. Jalapeño cannot train. Narrowing the job is the whole design bet, and perf-per-watt is where that bet either pays or does not.
When Does Jalapeño Get Deployed?
OpenAI says it plans to begin deploying Jalapeño across its own compute infrastructure by year-end 2026. The company also says a second generation is deep in development and a third is taking shape — which reads less like a one-off experiment and more like a standing silicon programme.
The broader signal for readers following our AI coverage is that custom inference silicon has moved from announcement to measurement. The next milestone worth waiting for is independent benchmarking outside the vendor's own lab.
Sources: OpenAI — August 25, 2026; TechCrunch — August 25, 2026; Tom's Hardware — August 26, 2026.
More AI Stories

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

Model Hardware Standard Lets AI Agents Run Lab Gear
Anthropic's Model Hardware Standard cuts lab instrument integration from weeks to minutes, with Genentech, Carnegie Mellon and QuEra among first users.

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.
