NVIDIA (NVDA-US) has recently revealed further technical details of its next-generation data center CPU, 'Rosa,' confirming it will adopt the new 'Rigel' CPU core architecture and continue with the Arm v9.2 architecture design, focusing on massive single-threaded computing power required for AI agent (Agentic AI) workloads. Rosa will be launched alongside the next-generation Feynman GPU platform, expected to enter the data center market in 2029, with a PC version, Spark, to be introduced in 2030.

NVIDIA first announced Rosa CPU at the 2026 GTC conference, stating it would be released together with the next-generation Feynman GPU as a new data center platform for the AI agent era. Named after American Nobel Physics laureate Rosalyn Sussman, Rosa succeeds the current Vera processor, further enhancing AI data center performance.

According to NVIDIA's latest information, Rosa will feature the new Rigel core, built on the Arm v9.2 architecture, and deliver improved single-core performance while maintaining the same chip area. Key improvements in Rigel include more efficient instruction delivery, larger L2 cache, and enhanced memory management capabilities to handle the heavy CPU workloads required by continuously running AI agents.

NVIDIA stated that the current Vera uses the in-house Olympus core, which achieves a 50% higher instructions per cycle (IPC) compared to Grace CPU and double the overall throughput. Vera features 88 Olympus cores, approximately 22% more than Grace's 72 cores, though NVIDIA has not yet disclosed whether Rosa will further increase the core count.

The company noted that Rigel will deliver higher single-threaded performance than Olympus while maintaining the same chip size, continuing NVIDIA's AI-centric CPU development strategy and directly competing with traditional x86 data center processors.

NVIDIA emphasized that as AI agents become an increasingly important computing paradigm in large-scale AI factories, the CPU's role is no longer limited to system control but now directly participates in the AI inference process, including tool calling, code execution, data processing, key-value cache (Key-Value Cache) management, and result validation.

Since GPUs are the most expensive and critical computing resources in AI data centers, insufficient CPU processing speed forces GPUs to wait, directly impacting the overall efficiency of the AI factory. Therefore, high single-threaded performance CPUs have become a new requirement for AI infrastructure.

NVIDIA pointed out that most current data center CPUs, driven by cloud computing demands, have shifted toward high core counts, sacrificing some single-core performance to reduce costs and increase computing density. However, this compromises cache capacity, memory architecture, and instruction processing capabilities, making them unable to meet the continuous, high-frequency, and low-latency demands of AI agent workloads.

For example, Vera uses a monolithic architecture, providing up to 1.2TB/s of LPDDR5X memory bandwidth, with memory power consumption below 40 watts and core-to-core interconnect bandwidth reaching 3.4TB/s—about three times higher than other data center CPUs—ensuring all 88 CPU cores maintain high performance simultaneously and avoid memory bottlenecks.

In AI agent workload tests, NVIDIA stated that Vera's sustained single-core performance reaches 1.8 times that of mainstream x86 server CPUs.

AI search startup Perplexity has already adopted Vera to test AI agent workflows. NVIDIA noted that in real-world code development scenarios, including copying codebases and running sandbox tests, Vera completes tasks about 1.5 times faster than x86 platforms, with multi-sandbox startup speeds improving by 1.9 times. Perplexity is currently planning to deploy the Vera platform in its production systems.

Additionally, tests with data analytics and real-time streaming platform partners showed a 3x speed improvement in large SQL analysis using Starburst and up to a 6x reduction in real-time streaming latency with Redpanda, both outperforming current mainstream x86 server platforms.

NVIDIA highlighted that Vera's greatest advantage lies in its ability to simultaneously handle diverse tasks such as AI agent tool execution, data analytics, code computation, reinforcement learning, and inference, eliminating the need for different CPUs for different workloads. It also serves as the CPU core for the Vera Rubin GPU platform and BlueField-4 STX storage processor, enabling the entire AI factory to adopt a unified architecture and development toolchain.

Looking ahead, NVIDIA plans to launch the RTX Spark platform, featuring Grace CPU and Blackwell GPU, this autumn; the Vera Rubin platform in 2028; the Rosa and Feynman data center platform in 2029; and the Rosa Feynman Spark PC version in 2030, gradually establishing a comprehensive product lineup spanning AI data centers and endpoint computing.

NVIDIA believes that in the future, tens of billions of AI agents will continuously perform inference, data retrieval, tool operations, validation, and decision-making, each highly dependent on CPU performance. As AI agents become the primary computing model in AI factories, faster CPUs allow GPUs to spend more time on high-value AI tasks, reduce waiting time, and increase overall data center productivity, with Rosa playing a crucial role in this strategy.

FACT BOX

  • Source: PR Times
  • Category: New Product
  • Organizations: Perplexity / Starburst / Redpanda
  • Products / services: Rosa CPU / Feynman GPU