OpenAI has unveiled its first custom silicon, an inference chip codenamed Jalapeño, and the early numbers are aimed squarely at Nvidia.

According to India TV News, the chip was developed with Broadcom and is designed primarily for inference — the stage where a trained model actually answers your questions, as opposed to the far more expensive training process that creates the model in the first place. Livemint makes the same distinction. Data Center Dynamics reports the chip carries a 700W TDP, and ServeTheHome says OpenAI gave a technical deep-dive on the design at the Hot Chips 2026 conference.

The performance claims come from OpenAI itself, which published results it describes as industry-leading for speed and efficiency in AI inference. As reported by The Verge, OpenAI says Jalapeño delivered 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower latency than Nvidia chips, measured across GPT-OSS, DeepSeek R1 and Kimi K2.5 1T. TechCrunch reports the testing used SemiAnalysis's InferenceX benchmark, where Jalapeño registered both more tokens per user and more throughput per kilowatt than the current state of the art. Neowin and Proactive frame the comparison against Nvidia's GB300, while SemiAnalysis headlines its own take as "Better Than Nvidia Blackwell."

Bloomberg and Axios both note these are OpenAI's own claims from its own tests — an important caveat until independent results arrive. The Register and Neowin also describe the chip as upcoming rather than shipping. One striking detail from The New Stack: OpenAI built the chip in nine months, then let AI rewrite the code.

Why it matters: if OpenAI's biggest customer-side cost — running models for hundreds of millions of users — can be served by its own cheaper, cooler silicon, the company loosens Nvidia's grip on the most profitable business in tech, a risk The Information suggests could "sicken" Nvidia.