The first thing OpenAI told the world about its own chip was how much electricity it uses. Not how fast it is, not what it costs, not how many it will build. Watts.
Jalapeno’s first results went up on August 25, timed to Hot Chips. The chip is rated at 700W and stayed at or under 550W in the tested workloads. Coverage of the presentation puts a rack of 128 at about 130kW. The performance claims are all denominated the same way: 1.5 to 1.9 times more work per watt than Nvidia’s GB200 and GB300 systems, on InferenceX runs of GPT-OSS 120B, DeepSeek R1 and Kimi K2.5.
— OpenAI, Jalapeno's first results, August 25 2026Although performance is sometimes reported per chip, we believe the more useful standard is performance per unit of power.
SemiAnalysis, whose team verified the runs in OpenAI’s lab, says why in one line: “OpenAI is currently limited by datacenter power, not by budget or floorspace,” so “tokens per MW is paramount.” Their gloss is blunter. “Throughput per watt is revenue.”
That sentence is worth more than the benchmark.
The Rack Draws the Same. The Rack Does More.
Run the arithmetic. 128 chips at 700W is about 90kW before networking and cooling, landing at 130kW per rack. Nvidia’s GB200 NVL72 rack, 72 GPUs at 1,200W each, lands in roughly the same place. The rack pulls the same megawatts. What changed is what it does with them. More, smaller packages, HBM4 stacked six deep, 15.4TB/s of memory bandwidth per chip, and no training hardware on the die. That is the honest description of the win, and it is real. It is also narrower than the coverage suggests.
What SemiAnalysis Actually Signed Off On
— SemiAnalysis, August 25 2026All numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results.
The caveats they list:
- One easy workload. 8,000 tokens in, 1,000 out. No long context, no agentic loop. AgentX, the benchmark for the work that earns money, has not been shown.
- Wrong comparison target. GB200 and GB300 are Blackwell. “Jalapeño is really competing against chips like Rubin,” and by their figures Vera Rubin NVL72 delivers 5.4 times the performance per megawatt of GB200 NVL72. Jalapeno’s runs used no speculative decoding. Rubin’s do. Against last year’s chip, Jalapeno wins by 1.5 to 1.9 times. Against this year’s, on the analysts’ own numbers, it may not win at all.
- Early silicon, no production realism. The tested part is the A0 stepping. Cache management and routing, the two things that decide real inference cost, were not exercised.
And the “50% cheaper inference” line in the headlines is Broadcom CEO Hock Tan, in a June interview, with no baseline. OpenAI’s own post never puts a dollar on anything. It talks about “operating leverage” and lets the reader do the rest.
Deployment starts at the end of 2026 in what OpenAI’s hardware lead calls very small volumes. Wider rollout is 2027. Gen 2 is “deep in development” and Gen 3 “taking shape.” By the time Jalapeno is a meaningful share of the fleet, these results will be about the previous chip. Nobody has said what share that will be, and the same post promises to keep deploying Nvidia widely.
Why Power Is the Price
Three days after the chip results I wrote about the week DeepSeek replaced flat token pricing with peak and off-peak rates, and OpenAI ran discounts that expired on a calendar. The sticker price on a token had stopped being information.
Jalapeno is the supply side of that story. If OpenAI’s binding constraint were capital, the fix would be more chips, and the price of a token would track the price of a chip. The constraint is power. So the fix is more tokens per megawatt, and the price of a token tracks the price of a megawatt, which has always been sold by the hour.
A company that optimises silicon for tokens per megawatt will optimise pricing for tokens per megawatt too. Off-peak discounts, three-month promotions, reseller-only rates. Those are not marketing experiments. They are the shape of a product whose input is electricity.
Jensen Huang’s response on CNBC was “lots of projects get started, lots of projects get canceled,” and he is entitled to it. Jalapeno does not touch training, where Nvidia’s margin lives. Broadcom’s stock still rose 4.5% on the week. The market is not pricing a Nvidia killer. It is pricing the largest inference customer on earth telling everyone which metric it cares about.
What This Does Not Settle
- Whether the advantage survives Rubin. A fair comparison needs Rubin, speculative decoding on, at long context. Nobody has published one, and the one number available points the wrong way.
- Whether tokens per megawatt is the right metric for agents. Agentic workloads are dominated by prefill on long contexts and cache hits. A chip tuned for an 8k/1k benchmark can look very different on a 200k-token coding session.
The Takeaway
Discount the multiples. Keep the sentence. OpenAI let the analysts it invited into the lab say, on its behalf, that the thing it runs out of is not money. When the richest lab in the industry says its ceiling is megawatts, every pricing decision it makes afterwards should be read as a power decision wearing a token’s clothes.
The chip is inference-only, early silicon, measured on the easy case against last year’s competitor. It still moves the argument, because it tells you what OpenAI thinks the argument is about. It was never about the chip.



