Nvidia makes the buzz rolled out its new Vera Rubin chip system this week, revealing new performance benchmarks for the GPU and CPU combination in front of his rival AMDSan Francisco’s annual product event on Thursday.
During a lengthy technical workshop last week at the company’s headquarters in Santa Clara, Calif., Nvidia executives bragged to a small group of journalists about the chip system’s increased power and efficiency capabilities. The biggest takeaway: Nvidia, which has long specialized in GPU manufacturing, is increasingly trying to position itself as a supplier of processors capable of powering AI agents.
Even if GPUs remain the main material that companies use to train and run their AI models, the industry’s shift toward more complex agentic systems has increased demand for processors, capable of orchestrating data flows, networking, and other software tasks. This is probably one of the reasons why Nvidia wanted to present itself as a supplier of complete AI systems, rather than just AI chips.
Vera Rubin is Nvidia’s successor hybrid superchip Grace Blackwell system, and represents the lynchpin of its near-term future, powering the AI industry. It is designed to offer one CPU for two GPUs. In a single Vera Rubin NVL 72 superchip system, there are 36 Vera CPUs for 72 Rubin GPUs. Nvidia also sells the Vera processor as a standalone product, and it has reportedly told Chinese customers these could be ready as early as August.

The CPU chip that Nvidia uses for its new Vera Rubin hardware system.
Courtesy of NVIDIA
Nvidia executives emphasized that its new Vera Rubin NVL72 racks – a stack of chips packed into a single liquid-cooled platform – are much more “plug-and-play” than some of its previous products. During a brief tour of an Nvidia data center lab in Silicon Valley, Nvidia executives said OpenAI already uses a Vera Rubin rack.
Nvidia CEO Jensen Huang didn’t make an appearance at the Santa Clara workshop last week; he was in Japan to announce the new partnerships with several Japanese companies develop AI for robotics. The briefings were led by Ian Buck, Nvidia’s longtime vice president of accelerated computing and the architect behind the company’s CUDA software.
“We are on a roadmap to create new architectures, not just GPUs but also CPUs,” Buck told reporters. “We’re going to continue to innovate, because it’s do this or die in Silicon Valley.”
The meetings took place in Huang’s executive briefing center, where several nearby desks were filled with bags of Taiwanese snacks that the CEO had brought back from his recent trip to Computex, a large annual semiconductor trade show in Taipei, an Nvidia spokesperson told WIRED.
A rack containing Nvidia’s Vera processors, designed to pair with the company’s next-generation Rubin GPUs in AI data centers.
Courtesy of NVIDIA
Nvidia claims that the Vera Rubin NVL72 system will process ten times more tokens per watt than the company’s Grace Blackwell superchip. The company says its Vera processor is also faster at processing agentic AI tasks compared to competing processors from AMD and Intel (although tests conducted to support these tests appear to have used slightly older generations of its competitors’ processors). The memory subsystems located on the new chips will also offer nearly three times more memory bandwidth than Blackwell’s, which will likely be an attractive feature for many companies amid a crisis. persistent shortage high bandwidth memory.
Nvidia says it has also significantly reduced the number of cables needed to connect its chips to the racks of multi-rack server systems, to the point where the company is billing Vera Rubin as “cableless compute” and “hot-swappable.” This means that customers can theoretically reduce the time it takes to install each rack from hours to minutes, a point made by Buck and Andrew Bell, Nvidia’s senior vice president of hardware engineering. And the new chip system is 100% liquid cooled, which can reduce the amount of energy needed to cool the chips, since air cooling is more energy intensive.
Since Nvidia unveiled Vera Rubin in spring 2025, the company has slowly released more details about the chip system while insisting on its on-time release. Huang has repeatedly said that Vera Rubin is reaching “full production” and will ship in the second half of this year, with early customers including Microsoft, OpenAI and Oracle.
Nvidia is particularly sensitive to any suggestions of delays after its previous generation Blackwell chips reportedly overheated when connected together in the company’s custom server racks, forcing it to make design changes and push back shipments.
Nvidia’s marketing push for Vera Rubin comes just before rival AMD’s annual conference, where executives are expected to tout its next-generation AI and data center chips. Sunday, AMD revealed more details about its Helios AI chip rack, designed to compete with Nvidia’s new products. AMD and Nvidia are arguing large-scale multi-year contracts to supply chips to AI hyperscalers like Meta and Amazon and AI labs like OpenAI, Anthropic and SpaceXAI.
Over the past two years, AMD has significantly increased its market share of processors used in data centers. The company has long been recognized as a pioneer of modern chip architecture used in x86 processors, which still account for the vast majority of data center processor revenue. Nvidia, on the other hand, builds its data center processors on ARM, an alternative chip architecture known for its power efficiency.
Buck and Hannah Coutand, product marketing managers for Nvidia DGX Cloud, both pointed out that Vera Rubin is abandoning the chiplet architecture used by many modern processors in favor of a single monolithic chip. Coutand argued that assembling multiple chipsets imposes “a heavy tax on memory bandwidth and data movement”, whereas Vera Rubin’s monolithic design allows data to move more quickly on a single integrated circuit.
