Nvidia is preparing the next big leap in the hardware that powers artificial intelligence, and the headline claim is about money: cheaper AI.
According to Tech Times, Nvidia's new Vera Rubin platform is set to ship this fall, with eight cloud partners lined up to deploy it. The report says the system is designed to deliver a 10x lower "token" cost — meaning the price of generating each chunk of AI output, the work known as inference, could drop sharply compared with today's hardware.
A key piece of that improvement is memory. Tech Times reports that Vera Rubin uses HBM4, a new generation of high-bandwidth memory that triples the rate at which data moves to and from the chip. Faster memory matters because modern AI models are often bottlenecked not by raw computing power but by how quickly they can feed data through — so tripling bandwidth directly helps the chip do more work per second.
Vera Rubin is not a standalone product. According to semivision, Nvidia is rolling out six chips aimed at the next generation of AI data centers in 2026, signaling a broad refresh of the company's lineup rather than a single new part.
Why it matters: the cost of running AI — not just training it — has become one of the biggest constraints on how widely the technology can be used, and if Nvidia's claimed 10x reduction in token cost holds up in practice, it could make AI features dramatically cheaper to offer at scale.