NVIDIA (NVDA-US) is expanding the artificial intelligence (AI) chip competition from the GPU market into the CPU domain. On May 21 Eastern Time, just before AMD's AI-related event, NVIDIA announced the latest progress on its next-generation Vera Rubin platform, revealing that the Vera Rubin NVL72 has entered the mass production ramp-up phase. The supply chain now spans over 350 factories across 30 countries, with more than 300 partner companies involved in deployment.

Simultaneously, NVIDIA released new performance data for its Vera CPU and Vera Rubin platform on AI agent workloads, aiming to demonstrate that as AI applications evolve from simply 'answering questions' to autonomous planning, tool invocation, and task execution, CPUs are becoming a critical battleground in AI infrastructure.

NVIDIA stated that full specifications, benchmark results, and architectural details of its data center CPU product Vera are essential for potential customers evaluating the chip. Vera chips have already been delivered to customers including OpenAI, Anthropic, and SpaceX since June.

The launch of Vera marks a key step in NVIDIA's vertical integration strategy. The company aims to provide customers with a complete AI computing platform encompassing its own CPUs, GPUs, networking chips, and software systems, rather than selling individual chips.

Research firm Wolfe Research estimates that the average selling price of a Vera chip is around $5,000, with an expected shipment volume of approximately 1.3 million units this year. If this forecast materializes, NVIDIA will formally enter the server CPU market long dominated by Intel (INTC-US) and AMD (AMD-US), directly challenging the core businesses of both companies.

NVIDIA announced in late May that Vera had entered full production, noting that the chip achieves up to 1.8 times faster task completion speed than traditional x86 CPUs in specific workloads. In mid-May, the company also delivered initial Vera CPU systems to customers such as Anthropic, OpenAI, and SpaceX AI.

The focus of this latest announcement shifts to the mass production progress of the overall Vera Rubin platform, customer deployment status, and real-world performance in production environments.

NVIDIA's strategic bet on CPUs stems from the rapid transformation of AI application patterns. Previously, generative AI primarily relied on GPUs for large-scale model training and inference, while CPUs handled peripheral tasks such as data processing and system scheduling. However, with the rapid development of AI agents, AI systems are now autonomously decomposing tasks, invoking external tools, executing programs, accessing data, and repeatedly evaluating results—elevating the importance of CPUs.

NVIDIA argues that AI agents are not workloads solely dependent on GPUs. Each agent's execution environment, tool invocation, task orchestration, and long-context data retrieval require CPUs to deliver low-latency, high-efficiency processing.

When NVIDIA first introduced Vera, it stated that agent-based AI is creating a new 'CPU moment.' As AI systems transition from 'answering questions' to 'taking actions,' the role of CPUs within the AI data center will significantly increase.

NVIDIA even forecasts that the long-term market size for server CPUs could reach $200 billion. This is an aggressive projection compared to the traditional server CPU market, but NVIDIA believes the boundaries of the CPU market will expand beyond enterprise servers into new domains such as AI inference, AI agents, reinforcement learning, data processing, and AI infrastructure control.

Vera's product design differs from traditional server CPUs. NVIDIA states that Vera is its first CPU purpose-built for agent-based AI, featuring 88 custom-designed Olympus cores and 1.2TB/s of memory bandwidth.

Compared to traditional CPUs, Vera achieves a 50% improvement in single-core performance. In prior tests, Vera demonstrated 1.8 times faster task completion speed than x86 CPUs.

In this latest update, NVIDIA further emphasizes that Vera's design philosophy is not merely stacking more cores, but enhancing single-thread performance, inter-core communication bandwidth, and memory access efficiency.

The custom Olympus cores used in Vera deliver 2x single-thread performance, 3x inter-core bandwidth, and reduce memory latency by 40% compared to competitors' chiplet designs.

This enables AI agents to complete tasks faster and return computing resources to GPUs more quickly, improving overall AI data center utilization efficiency.

Since GPUs are currently among the most expensive and scarce computing resources in AI data centers, insufficient CPU processing speed can cause GPUs to idle while waiting for data, reducing utilization. NVIDIA aims to minimize such 'idle time' through specialized CPU architecture optimization.

In production environment testing by DeepInfra, NVIDIA claims Vera supports up to 1.6x more concurrent AI agents and achieves up to 2.2x faster task orchestration speed. However, NVIDIA notes that these figures come from specific customer test environments and do not indicate Vera's universal superiority across all general-purpose CPU workloads.

The Vera CPU is only part of NVIDIA's broader AI system strategy. Within the Vera Rubin platform, NVIDIA integrates Vera CPU with Rubin GPU, Groq 3 LPX, Spectrum-6 networking chip, ConnectX-9 SuperNIC, and BlueField-4 to build a complete AI infrastructure.

The Vera Rubin NVL72 is currently ramping up mass production globally, with partners such as CoreWeave, Google Cloud, Microsoft (Microsoft, MSFT-US) Azure, and Oracle (Oracle, ORCL-US) Cloud Infrastructure beginning to deploy the rack-scale systems.

NVIDIA reports that the entire supply chain now covers over 350 factories and 30 countries, with more than 300 partner companies involved.

In customer testing, CoreWeave reported that running the DeepSeek-R1 model on Vera Rubin NVL72 achieved a 10x increase in tokens generated per second per terawatt compared to the previous-generation Grace Blackwell NVL72.

Google Cloud has already launched A5X cloud instances based on Vera Rubin NVL72. NVIDIA states that in specific application scenarios, this platform delivers lower inference costs and higher token throughput per terawatt.

This indicates that NVIDIA's competitive target is no longer limited to single AI accelerator chips like AMD Instinct GPU or Intel Gaudi, but aims to control the entire AI data center architecture.

In NVIDIA's vision, CPUs handle task orchestration, GPUs perform high-intensity computing, networking chips enable high-speed interconnects, and software manages resource optimization—ultimately delivering rack-level or even data center-level solutions to customers.

For AMD and Intel, the biggest challenge is that NVIDIA does not need to engage in direct competition in the traditional CPU market, but instead targets the fastest-growing AI server use cases such as agent-based AI, high-performance inference, and reinforcement learning.

While Intel and AMD have long dominated the server CPU market, NVIDIA believes AI infrastructure is reshaping the value hierarchy of CPUs.

Historically, server CPU competition centered on core count, general-purpose computing power, and overall throughput. However, in the era of agent-based AI, low latency, single-core performance, memory bandwidth, and task orchestration efficiency are gaining prominence.

This is precisely the new market opportunity NVIDIA believes it can exploit.

In terms of product positioning, Vera can be used as a standalone CPU or combined with NVIDIA GPUs to form a complete Vera Rubin system. NVIDIA positions it as the critical core connecting AI agents with GPU computing resources.

Previously announced customers include Anthropic and OpenAI.

FACT BOX

  • Source: PR Times
  • Category: New Product
  • Organizations: AMD / Intel / OpenAI
  • Products / services: Vera CPU / Vera Rubin NVL72