The technology industry is shifting its focus from GPU shortages to CPU, a foundational computing resource.
According to The Information, Amazon (AMZN-US) cloud service (AWS) issued internal warnings to its engineering teams in May this year, urging them to conserve computing resources to ensure continued ability to meet customer demand.
Notably, this conservation directive extends beyond AI chips to traditional CPU servers that have long supported web services.
Insiders reveal AWS has set deadlines for departments to reduce their computing resource usage by later this year. To free up processing power for external customers, engineers are progressively shutting down EC2 virtual servers previously used for software development that are now idle.
A long-time AWS engineer reported that while CPU server resource requests were typically fulfilled within hours in the past, they now face delays of several days—an unprecedented situation in his career that could potentially delay project deliveries.
In response, AWS officially stated that despite currently facing "massive demand," the company remains capable of meeting the computing needs of "the vast majority" of internal and external customers. The company emphasized it is working closely with internal teams to maintain efficient EC2 resource utilization, and employee resource usage guidelines have not changed due to the current tight situation.
Currently, the impact on external customers remains limited. A consultant assisting enterprises with AWS integration noted no observed shortages in contractually committed computing power.
However, Spot Instances—offered at discounted prices but subject to termination within two minutes—have become increasingly difficult to obtain in large quantities over recent months. This signals a narrowing supply-demand gap across AWS and shrinking flexibility, potentially foreshadowing rising cloud computing costs.
Analysis indicates the fundamental cause of this CPU crunch is the rapid proliferation of "AI agent" applications. Jing Xie, co-founder and managing director of AI consultancy Elendil Labs, observed that per-customer IT spending has doubled, primarily due to enterprises widely adopting AI agents for software development.
He noted that AI agents require significantly more cloud resources during operation, creating sustained heavy demand on CPUs.
"Regardless of current development, production, or operations, almost every task now requires more CPU than before," he stated.
In fact, CPUs play roles beyond executing AI applications—they are also essential in the AI development process. For example, AI companies still rely on CPUs to read and process raw data from documents, images, and videos when preparing data for model training.
Recent statements from major chipmakers confirm this trend. Intel (INTC-US) CEO Lip-Bu Tan revealed during the April earnings call that in AI inference (actual model operation) scenarios, the CPU-to-GPU usage ratio was approximately 1:4.
However, by July, the company's CFO David Zinsner indicated this ratio had rapidly narrowed to nearly 1:1.
Senior executives from AMD (AMD-US) and Arm (ARM-US) have made similar statements, indicating CPU demand growth has exceeded industry expectations.
Beyond CPU supply shortages, tight memory chip supply for accompanying use and physical data center space limitations are collectively exacerbating this computing power shortage.
Large Tech Companies' Computing Resource Allocation Challenges
Allocating limited computing power between internal R&D and external customers has become a shared challenge for tech giants over the past two years.
According to previous reporting by The Information, Google established a dedicated committee of senior executives last year to coordinate computing resource allocation among Google Cloud, DeepMind research division, and consumer-facing businesses.
Despite this, internal friction persists. Google's renowned AI researcher Noam Shazeer chose to leave earlier this year due to dissatisfaction with restricted computing access.
In contrast, Microsoft has demonstrated positive benefits from improved computing management efficiency. The company's CFO Amy Hood stated last month during the earnings call that part of Azure cloud service's revenue growth was precisely due to improved management efficiency of CPU and GPU clusters.
FACT BOX
- Source: PR Times
- Category: News
- Organizations: Amazon / Google / Microsoft
- Products / services: EC2 / Spot Instances