OpenAI has spent years buying enormous amounts of computing power from other companies. Now it has a chip of its own, and the first serious benchmark results suggest the newcomer is not merely a bargaining chip with a cute name.
Jalapeño, OpenAI's first custom inference accelerator, delivered more AI work per unit of power and lower latency than leading commercially available systems in tests released Tuesday. The chip was developed with Broadcom and is designed specifically for inference, the part of AI computing that happens after a model has been trained and is responding to users.
That distinction matters. Training gets the spectacular data center photos and the billion dollar procurement headlines, but inference is where ChatGPT, coding agents, APIs, and other AI products burn through compute every minute of the day. As usage grows, even small improvements in efficiency can turn into very large differences in cost.
The benchmark numbers are hard to ignore
OpenAI tested Jalapeño using InferenceX, the public benchmarking platform built by SemiAnalysis. Across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, OpenAI says Jalapeño landed on the performance frontier for both speed and efficiency.
On the largest model tested, Kimi K2.5, Jalapeño delivered about 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. Across the broader test set, reported gains reached roughly 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency.
SemiAnalysis, which says it was invited into OpenAI's labs to examine and benchmark the hardware with its InferenceX suite, described the first generation chip as unusually competitive and said it beat Nvidia, AMD, and Google hardware the firm had tested on several major open models.
That is a notable result because first generation custom accelerators often spend their debut proving that they work at all. Jalapeño appears to have skipped that awkward phase and gone straight into a comparison with Nvidia's Blackwell generation.
OpenAI designed it around one job
The reason Jalapeño can look so strong against more general accelerators is also the reason the comparison needs context. Nvidia's GPUs are built to support a broad range of workloads, including AI training. Jalapeño is an application-specific integrated circuit built around serving large language models.
OpenAI says it designed the chip, memory system, networking, software, and serving stack together so less time and energy are lost moving data around. The accelerator is rated at 700 watts, while OpenAI says sustained measured power stayed at or below 550 watts on the workloads it tested.
The company also used its own AI models during the design process. When Jalapeño was first unveiled in June, OpenAI said the core design-to-tape-out process took nine months, with Broadcom handling silicon implementation and networking technology and Celestica helping with boards, racks, and system integration.
OpenAI's detailed benchmark report says internal testing showed an even larger advantage on frontier OpenAI models, although those results are not independently comparable because the models are not publicly available.
This is about cost and leverage as much as speed
OpenAI is not planning to sell Jalapeño as a standalone commercial chip. Its value is internal. Every query that can be served more efficiently is a query that costs OpenAI less money, uses less power, and requires less third-party hardware capacity.
That gives the company another kind of leverage too. Nvidia remains central to OpenAI's infrastructure, and OpenAI says Jalapeño is meant to sit alongside hardware from Nvidia, AMD, Microsoft, AWS, Cerebras, and other partners rather than replace all of it. Still, becoming one of your supplier's largest customers is good. Becoming a customer that can also build credible alternatives is considerably better.
The timing matters because AI companies are increasingly constrained by electricity, data center capacity, and the price of high-end accelerators. Faster inference is useful, but higher throughput per kilowatt may be even more important as new data centers compete for limited power.
Nvidia will not be standing still
There is one important reason not to declare a new chip king after a single benchmark cycle. Jalapeño is only beginning limited deployment in 2026, with broader production expected to ramp in 2027. Nvidia's roadmap will keep moving during that same period.
Today's comparison is primarily against Blackwell generation hardware that is commercially available now. By the time OpenAI deploys Jalapeño at much larger scale, Nvidia will have newer systems in the field. Custom chips also have to prove more than benchmark speed. Yield, reliability, software maturity, networking performance, and the ability to run at data center scale all matter.
Still, OpenAI does not need Jalapeño to replace Nvidia to make the project worthwhile. If its own silicon can take a meaningful share of inference traffic while cutting power and latency, it changes the economics of serving AI and gives OpenAI more control over one of its biggest costs.
The first numbers suggest Jalapeño is already more than an experiment. OpenAI now has working first-party silicon that can compete with established AI hardware on the workload it cares about most. For a company whose products are increasingly limited by how cheaply and quickly they can answer the next billion prompts, that could matter just as much as another leap in model intelligence.
---------------
Author: Ryan Gardner
Silicon Valley News Desk