NVIDIA (NVDA-US) and AMD (AMD-US) are not merely competing over whose CPU is faster—they are vying to define the very standard by which CPUs are evaluated in the AI era.
According to a report by Wall Street Horizon, NVIDIA recently disclosed the most detailed technical specifications of its Vera CPU architecture. This processor features 88 proprietary Olympus ARM architecture cores, 1.2TB/s of memory bandwidth, and 3.4TB/s of on-die interconnect bandwidth. NVIDIA's core argument is that AI agents operate through repeated, sequential interactions between CPU and GPU—tool invocation, code execution, data retrieval, task orchestration—each step dependent on the completion of the prior one. Therefore, the single-core speed and latency of the CPU directly determine the overall response efficiency of the agent.
On July 23, according to Feng News, Bank of America Securities analyst Vivek Arya stated in a new report that Vera's launch introduces a novel evaluation framework to the industry: 'maximum single-threaded performance at scale.' This stands in direct contrast to AMD's long-standing chiplet-based multi-core stacking approach. The firm frames this debate as: 'Is Agent AI bottlenecked by the time to complete a single task, or by how many concurrent tasks a single rack can run?' AMD is set to hold its AI 2026 Technology Day this Thursday, which will be its first formal public response to this framework.
This debate unfolds against the backdrop of an AI-driven transformation in the server CPU market, projected to grow nearly fourfold by 2030 to $170 billion. This is not a zero-sum game, but an emerging incremental market—whichever company establishes an industry-recognized performance benchmark first will gain pricing power and narrative control.
NVIDIA describes the AI agent workflow concretely: during task execution, the CPU and GPU engage in numerous serial interactions, each call waiting for the prior step to complete. This means any delay in any stage accumulates and amplifies, ultimately slowing down the entire AI factory's output efficiency.
Under this logic, single-core performance is not just a spec-sheet number—it is a critical variable directly impacting GPU utilization. The faster the CPU, the less time the GPU spends waiting, and the higher the overall system resource efficiency. In short, the faster the single core, the quicker the entire pipeline responds.
Another notable design choice in Vera is NVIDIA's decision to use a monolithic die instead of a chiplet-based disaggregated architecture, citing better 'scalable coherency.' This contrasts with AMD's long-term bet on chiplets, reflecting a deep philosophical divergence in underlying architectural approaches.
Moreover, Vera is not a standalone product but part of NVIDIA's co-designed AI infrastructure ecosystem, working in tandem with Rubin GPUs, Groq LPX, Spectrum switches, and BlueField storage/NICs. This system-level integration is NVIDIA's true moat beyond mere hardware comparisons.
From AMD's perspective, production-grade AI systems are not about a single agent serially progressing through tasks, but more akin to a distributed software platform composed of databases, APIs, vector stores, orchestration engines, caches, and middleware. In this scenario, the bottleneck is not the speed of completing a single task, but how many workflows can be concurrently hosted within a fixed power budget.
AMD has provided specific calculations: in a simulated 100-kilowatt rack deployment scenario, the rack-level throughput of the EPYC 9965 (Turin) is approximately 2.4 times that of NVIDIA's Vera baseline; the next-generation EPYC 6 (Venice) is expected to reach 3.3 times.
AMD's logic is clear: higher throughput density means more users served and more requests processed under the same energy consumption—this, they argue, is the true cost function for cloud AI deployment.
The CPU architecture battle extends even to the instruction set level.
NVIDIA's choice of the ARM architecture implies the belief that as long as the microarchitecture is superior, instruction set compatibility can take a backseat. However, AMD and Intel are likely to continue emphasizing the opposite: AI is expanding from model inference into enterprise software workflows, and enterprise software stacks—across databases, middleware, security platforms, and enterprise applications—have accumulated decades of optimization, validation, and compatibility on the x86 ecosystem.
This is not merely a technical debate. As AI workloads become increasingly embedded into existing enterprise IT systems, the historical accumulation of software ecosystems may prove harder to displace than peak hardware performance. NVIDIA's ability to penetrate these domains with Vera will largely depend on the maturity speed of the ARM software ecosystem.
Who will define the next CPU KPI?
Two frameworks, two KPIs—this is fundamentally a battle for narrative dominance in the industry.
NVIDIA's framework is built around latency, single-thread progress, and GPU utilization. AMD's framework centers on concurrency, throughput, and service density.
The key question is: which metric will the market ultimately use to purchase server CPUs?
Bank of America's assessment is that AMD's Thursday event will not focus on outperforming rivals in benchmark scores, but on persuading the industry to adopt its evaluation system. Whichever framework is adopted by data center buyers will gain pricing power in this $170 billion market.
Bank of America maintains a 'Buy' rating on NVIDIA stock, with a target price of $350.
FACT BOX
- Source: PR Times
- Category: New Product
- Organizations: AMD / Bank of America Securities
- Products / services: Vera CPU / Rubin GPU