According to The Information, Alphabet's Google is developing a next-generation AI server chip codenamed 'Frozen v2,' designed to hardwire the underlying architecture of its Gemini large language model directly into silicon to enhance AI inference efficiency. Following the news, Alphabet (GOOGL-US) shares rose nearly 3% during Monday’s trading session, reflecting market optimism about its continued investment in AI infrastructure.
The report highlights that Frozen v2’s key innovation lies in embedding parts of Gemini’s architecture directly into the chip’s hardware, rather than relying solely on software computation. Compared to widely used general-purpose AI chips like NVIDIA’s GPUs and Google’s own TPUs, this new design trades some versatility for significantly higher inference efficiency and lower energy consumption.
Insiders revealed that Google internally estimates Frozen v2 could process 6 to 10 times more tokens per unit of power than the latest TPU generation. This efficiency gain stems from drastically reducing the logical decision-making and data movement required during model execution, enabling certain computations to be handled directly by hardware.
Current AI chips, such as NVIDIA’s GPUs and Google’s TPUs, use general architectures capable of supporting various large language models. However, each inference run still requires real-time data transfer and computation. In contrast, Frozen v2 adopts a highly customized design, permanently etching parts of Gemini’s base architecture into the chip, prioritizing 'lower flexibility, higher efficiency.'
A major driver behind this project is Google’s growing AI compute shortage. The report notes that internal GPU and TPU resources are strained, already affecting some Google Cloud customer services and forcing the company to decline certain external collaboration requests.
The report also states that Google recently agreed to pay SpaceX nearly $1 billion per month to access AI compute power from xAI, aiming to bridge its infrastructure gap. This highlights how rapidly growing AI inference demand is placing immense pressure on major tech companies’ computing resources.
In fact, the Frozen project is not new. An early version was led by Google DeepMind’s Chief Scientist Jeff Dean, originally planning to burn full model weights into the chip. However, this approach was shelved because it would only support a single Gemini version—rendering the hardware obsolete upon model updates.
The new Frozen v2 adopts an 'elastic hardening' strategy, fixing only the model architecture, not the weights. This allows Gemini to be iteratively updated via weight changes, maintaining high inference efficiency while extending the chip’s usable lifespan.
Google currently positions Frozen v2 as a new technical branch outside its TPU product line, not as a TPU replacement. The company has no plans for large-scale production at this stage, with estimated output far below TPU volumes. Instead, it is seen more as an experimental platform to validate future highly specialized AI chip technologies.
However, Frozen v2 carries clear risks. Because the chip design heavily depends on Gemini’s current architecture, a major overhaul of Gemini’s base structure could quickly make the hardware incompatible—potentially resulting in 'obsolescence before mass production.' Thus, Google appears to be betting that Gemini’s architecture will remain stable long-term.
Across the industry, AI inference chips have become a new battleground for tech firms. Canadian startup Taalas employs a similar strategy to Frozen v2, hardwiring model logic into hardware, and has already raised over $200 million.
Additionally, companies like SambaNova, d-Matrix, OpenAI, and Microsoft (MSFT-US) are actively developing AI inference chips to improve energy efficiency and reduce inference costs, challenging NVIDIA’s GPU dominance in the AI market.
Meanwhile, NVIDIA continues to strengthen its position. The report notes that in December 2025, NVIDIA spent $20 billion to acquire technology licenses from Groq, an AI inference chip company—indicating the global AI chip race is shifting from training compute to inference efficiency.
Market analysts believe that as large language models enter commercial deployment, the AI industry’s competitive focus is shifting from training capability to cost-effective, low-power large-scale inference. If Frozen v2 achieves its internal target of 6 to 10 times efficiency gains, it could reshape AI inference cost structures and further strengthen Google’s vertical integration advantage across models, chips, and cloud services.
However, analysts caution that Frozen v2 may not be deployed before 2028, as it remains in the R&D phase with a long validation process ahead for tape-out, testing, and mass production. Google currently has no plans for large-scale production, so its short-term impact on revenue and cost structure is limited. Moreover, the projected 6 to 10 times efficiency gain remains an internal engineering estimate, not yet validated by actual products.
Therefore, the market widely views the recent Alphabet stock rise as reflecting investor confidence in Google’s ongoing AI innovation, rather than a belief that Frozen v2 already possesses immediate, transformative value for the company’s profitability.
FACT BOX
- Source: PR Times
- Category: New Product
- Organizations: Alphabet / Google / NVIDIA
- Products / services: Frozen v2 / Gemini