The next frontier in AI computing is no longer just GPU speed, but how many tokens can be generated per watt of power. NVIDIA (NASDAQ: NVDA) announced today (22nd) that its next-generation Vera Rubin NVL72 is entering mass production, with CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure already running the system racks.

According to initial real-world testing by CoreWeave on the DeepSeek-R1 model, the Vera Rubin NVL72 produces 10 times more tokens per million watts per second compared to the Grace Blackwell NVL72. Google Cloud also claims that its next-generation A5X instances, powered by Vera Rubin, can reduce inference cost per token to as low as one-tenth of the previous generation.

As AI evolves from chatbots to agent-type AI capable of reasoning, planning, tool usage, and task execution, token consumption could reach 15 times that of traditional AI applications. With data center power supply increasingly constrained, the ability to generate more tokens with the same power—and reduce cost per million tokens—has become critical for cloud providers and AI developers to scale commercially.

CoreWeave (NASDAQ: CRWV), after months of co-development, became the first AI cloud provider to deploy and validate the Vera Rubin NVL72, releasing the first performance data measured on actual hardware. In benchmarking DeepSeek-R1 on Vera Rubin NVL72, CoreWeave found a 10x increase in tokens per million watts per second versus Grace Blackwell NVL72. This figure does not mean all workloads on Vera Rubin see a 10x performance gain, but rather reflects system efficiency under specific model and power conditions—how many tokens can be produced within a given power budget.

DeepSeek-R1 uses a Mixture-of-Experts (MoE) architecture, where each token must be routed across distributed expert sub-networks, making GPU-to-GPU communication a direct bottleneck. Vera Rubin NVL72 leverages a full NVLink 6 interconnect architecture with a total bandwidth of 260 TB/s, enabling the entire rack to operate like a single large accelerator and minimizing data transfer bottlenecks.

For AI cloud providers, producing more tokens per million watts means they can handle more workloads within the same power capacity—or reduce power demand and service costs if workloads remain constant. Enterprises and AI labs, including Jane Street, will access the Vera Rubin platform via CoreWeave’s cloud.

Vera Rubin’s performance gains aren’t solely due to the Rubin GPU. NVIDIA co-designed the entire platform—from chips and rack trays to networking and cooling—integrating seven chips and five rack tray types, including Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX.

NVIDIA states these components were designed from the start as a unified AI system, not as separate off-the-shelf parts later optimized. This strategy reflects NVIDIA’s evolution from a GPU supplier to a full-stack AI infrastructure platform encompassing CPU, GPU, networking, DPU, software, and cooling systems.

The Vera CPU is a processor specifically designed for agent-type workloads. According to NVIDIA’s data, its custom Olympus core delivers 2x higher single-thread performance, 3x higher inter-core bandwidth, and 40% lower memory latency compared to competitors’ chiplet designs. AI cloud platform DeepInfra further tested that, under identical service quality, Vera CPU can support up to 1.6x more concurrent AI agents, with coordination speeds up to 2.2x faster than other CPUs. DeepInfra currently processes nearly 5 trillion tokens weekly, with about 30% from agent-type systems—highlighting the CPU’s critical role in model invocation, data movement, and tool usage.

Vera Rubin has also entered Google Cloud. Alphabet (NASDAQ: GOOGL), Google’s parent company, launched the A5X bare-metal instance based on the Vera Rubin NVL72 rack-scale system, combining it with ConnectX-9 SuperNIC and Google’s next-gen Virgo networking.

According to Google Cloud’s roadmap, A5X clusters can scale to tens of thousands of Rubin GPUs at a single site, and multi-site configurations can approach nearly one million GPUs—used for training, fine-tuning, and deploying frontier models, open models, agent AI, and physics AI models.

London-based startup Ineffable Intelligence has already adopted A5X to develop a 'super learner' system that continuously learns from environmental interaction and experience. Unlike LLMs relying on static datasets, this system uses reinforcement learning across massive parallel simulations to generate experience, then converts results into strategy updates and evaluations—requiring high compute power, memory bandwidth, and low-latency interconnects.

Lasse Espeholt, co-founder of Ineffable Intelligence, said: 'Next-generation research needs next-generation hardware.' He noted the company completed Vera Rubin system activation almost immediately and has begun testing infrastructure for super learners.

Microsoft (NASDAQ: MSFT) and French AI startup Mistral will also expand their European AI infrastructure collaboration based on Vera Rubin. Under a new multi-billion-dollar agreement, Mistral will deploy thousands of Vera Rubin GPUs to expand compute resources for model training, inference, and large-scale deployment.

Mistral Medium 3.5 and OCR 4 are now in Microsoft Foundry, and Mistral models are integrated into Microsoft Copilot Studio. Through Azure Local and Foundry Local, government agencies and regulated enterprises can run the same models in public cloud, cloud-connected private environments, or fully offline settings.

This deployment reflects Europe’s unique AI market needs. Financial, healthcare, and government institutions demand not only compute performance but also data control, local deployment, regulatory compliance, and infrastructure resilience. Thus, Vera Rubin is not just a tool for Microsoft and Mistral to scale compute, but a foundation for entering Europe’s sovereign AI market.

To address rising AI system power and cooling demands, Vera Rubin NVL72 raises liquid cooling inlet temperature to 45°C, enabling operation with dry cooling and closed-loop liquid cooling systems without chillers. NVIDIA estimates that for new AI factories deploying 1 million kW capacity, this could save millions of gallons of water annually.

After three generations of rack-scale co-design, Vera Rubin’s compute trays no longer contain cables, fans, or hoses—reducing tray assembly time from hours to about one minute.

From CoreWeave’s real-world tests, Google Cloud’s near-million-GPU multi-site planning, to Microsoft and Mistral’s European deployment, it’s clear that Vera Rubin’s competitive focus has shifted beyond single-GPU performance. As the AI industry faces constraints in power, cooling, and inference costs, the ability to produce more tokens with less energy will determine the economic viability of next-generation AI infrastructure.

FACT BOX

  • Source: PR Times
  • Category: New Product
  • Organizations: CoreWeave / Google Cloud / Microsoft Azure
  • Products / services: Vera Rubin NVL72 / NVLink 6