OpenAI has unveiled Jalapeño, its first custom AI inference chip, and says early benchmark results show it running AI models faster while drawing less power than the hardware the industry currently relies on.

The pitch from CEO Sam Altman, widely quoted across coverage of the announcement, was blunt: "We made a chip and it is fast."

The headline claim is a direct challenge to Nvidia. According to reporting from The Information, OpenAI says Jalapeño outperforms Nvidia's Blackwell generation; other outlets, including Digit and Business Standard, report the comparison was specifically against Nvidia's current GB300 systems in key inference tests. The Free Press Journal reports the chip was tested across three open models — GPT-OSS 120B, DeepSeek R1 and Kimi K2 — with gains in speed, latency and power efficiency.

A note on what "inference" means here: training is the expensive process of building a model, while inference is the everyday work of actually answering your questions. Inference is where the electricity bill and the wait time live, so a chip that delivers more work per watt goes straight to OpenAI's costs and your response times.

OpenAI told reporters the chip was designed specifically around modern and future language-model workloads, with the processor, memory, networking, software and rack-scale system built together rather than assembled from off-the-shelf parts, according to Republic World. TrendForce reports Samsung is supplying HBM4 memory for the chip.

On timing, CNBC-TV18 reports OpenAI plans to begin using the chips to support its models later this year, while Firstpost reports wider deployment is planned for 2027. OpenAI has said it will run its own silicon alongside Nvidia and other accelerators, not instead of them.

It matters because OpenAI's biggest supplier is now also its rival, and if the claims hold up, the company that defines AI's software could start controlling the cost of running it.