AlphaWire

newswire

Google develops a new "Frozen" chip to significantly improve the energy efficiency of AI model operation

2026-07-20·newswire-us-stock-182737
Google develops a new "Frozen" chip to significantly improve the energy efficiency of AI model operation.

Two people familiar with the matter revealed, A new server chip is being developed. This chip can directly solidify the underlying architecture solution of the Gemini large model into silicon-based hardware, which can greatly improve the operating efficiency of Google's AI services for users.

This chip, internally codenamed Frozen v2 (Second Generation), is used to alleviate Google’s current severe AI computing power gap. Insufficient computing power not only caused conflicts among various departments within Google, but also caused Google Cloud to have to reject a large number of computing power purchase orders from external customers.

Engineers involved in the research and development of the chip calculated that based on the number of tokens that can be processed per unit power consumption (the core measurement indicator of AI computing power), after the chip is officially launched, the energy efficiency can reach 6 to 10 times that of Google's latest generation self-developed AI tensor processor (TPU).

The chip is named Frozen, derived from its design idea: part of the computing logic of the large model is permanently etched inside the silicon chip. Engineers are still finalizing the core functional modules of the chip and the collaborative working mechanism of each unit.

People familiar with the matter said that Google plans to officially put the chip into production as early as 2028. The Frozen project is not intended to replace Google's existing TPU tensor processors, but to create a new self-developed chip product line.

NVIDIA GPU and Google TPU, which are currently the most widely used in the industry, are general-purpose AI acceleration chips and are compatible with a variety of AI models.

When running different large models, general-purpose chips need to perform a large number of complex calculation scheduling decisions in real time; Frozen v2 builds some of the fixed calculation logic of the Gemini model into the hardware level, streamlining the chip's calculation steps and data transfer volume.

People familiar with the matter said that the hardware's built-in computing logic can shorten the response delay of AI question and answer and is expected to support Google's launch of a number of new AI applications.

But this plan also bets that Google will continue to use the research and development architecture of the current Gemini large model for a long time; only if subsequent iterations of the Gemini model are developed based on the underlying architecture of the chip design, can this chip be properly adapted to the new generation model.

The R&D personnel added that the chip retains flexible adjustment space and can complete model version upgrade iterations by updating the model weights (which determine the core parameter configuration of the model response logic).

Google’s official spokesperson responded: Google teams have been continuously developing and testing various innovative technologies, striving to deliver AI services with optimal performance and efficiency to users and enterprise customers.

The spokesperson also said: Not all R&D projects will eventually be put into mass production, but this kind of all-round technology exploration is a key part of Google's full-stack AI layout.

Why inference chips have become a hot spot in industry R&D A large number of companies are deploying AI inference chips (computing hardware that carries large models online), including startups such as SambaNova and d-Matrix, as well as OpenAI, And other technology giants.

The goal of each company is to create inference chips that are more energy-efficient than NVIDIA GPUs and alleviate the industry pain points of global computing power shortage and high computing power costs. Nvidia itself also spent $20 billion in December last year to acquire related technology licenses from inference chip startup Groq.

The technical routes of each manufacturer are different from Google's Frozen solution, but the design idea of Frozen v2 is similar to that of Canadian chip startup Taalas: both hard-code the computing logic of specific AI models into the chip hardware.

Taalas has completed over US$200 million in financing from investors including Quiet Capital and Fidelity Investments. The design of Frozen v2 is derived from the original Frozen solution, which was led by Jeff Dean, chief scientist of Google DeepMind, and planned to directly burn the model weights into the chip hardware.

However, this solution from many years ago had obvious shortcomings: the chip could only be adapted to a fixed version of the Gemini model, and the hardware service life cycle was too short, and it was eventually shelved.

Another person familiar with the matter revealed that the Google team is currently weighing the volume of solidified data built into the chip to find a balance between hardware energy efficiency and product versatility. Past earnings of Google’s self-developed chips Google’s years of self-developed chip layout has realized its value.

This year, Google officially released the eighth-generation TPU and began selling the chip products directly to cloud customers, directly impacting Nvidia's monopoly in the AI chip market.

Google has signed a multi-billion dollar TPU supply agreement with Meta, and is actively promoting its own TPU products to other cloud vendors that have been loyal to purchasing Nvidia hardware. Relying on self-developed chips and Gemini model co-design, Google's current hardware costs for running its own large models have dropped.

However, the existing TPU only embeds a small amount of model underlying information, and its integration level is far less than that of the Frozen series chips. Google currently has no plans to mass-produce Frozen v2 chips, and its production capacity is far lower than TPU.

The smaller mass production volume means that Google does not need to bring in external design and foundry partners early in the project.

Google also regards this second generation of Frozen as a technical experiment: after the future AI large model architecture becomes finalized, the experience accumulated in this research and development can help engineers design more specialized high-performance AI chips.

#Stocks #Nvidia #Meta #Google #AI

Full text

Google develops a new "Frozen" chip to significantly improve the energy efficiency of AI model operation

Two people familiar with the matter revealed, A new server chip is being developed. This chip can directly solidify the underlying architecture solution of the Gemini large model into silicon-based hardware, which can greatly improve the operating efficiency of Google's AI services for users. This chip, internally codenamed Frozen v2 (Second Generation), is used to alleviate Google’s current severe AI computing power gap. Insufficient computing power not only caused conflicts among various departments within Google, but also caused Google Cloud to have to reject a large number of computing power purchase orders from external customers. Engineers involved in the research and development of the chip calculated that based on the number of tokens that can be processed per unit power consumption (the core measurement indicator of AI computing power), after the chip is officially launched, the energy efficiency can reach 6 to 10 times that of Google's latest generation self-developed AI tensor processor (TPU). The chip is named Frozen, derived from its design idea: part of the computing logic of the large model is permanently etched inside the silicon chip. Engineers are still finalizing the core functional modules of the chip and the collaborative working mechanism of each unit. People familiar with the matter said that Google plans to officially put the chip into production as early as 2028. The Frozen project is not intended to replace Google's existing TPU tensor processors, but to create a new self-developed chip product line. NVIDIA GPU and Google TPU, which are currently the most widely used in the industry, are general-purpose AI acceleration chips and are compatible with a variety of AI models. When running different large models, general-purpose chips need to perform a large number of complex calculation scheduling decisions in real time; Frozen v2 builds some of the fixed calculation logic of the Gemini model into the hardware level, streamlining the chip's calculation steps and data transfer volume. People familiar with the matter said that the hardware's built-in computing logic can shorten the response delay of AI question and answer and is expected to support Google's launch of a number of new AI applications. But this plan also bets that Google will continue to use the research and development architecture of the current Gemini large model for a long time; only if subsequent iterations of the Gemini model are developed based on the underlying architecture of the chip design, can this chip be properly adapted to the new generation model. The R&D personnel added that the chip retains flexible adjustment space and can complete model version upgrade iterations by updating the model weights (which determine the core parameter configuration of the model response logic). Google’s official spokesperson responded: Google teams have been continuously developing and testing various innovative technologies, striving to deliver AI services with optimal performance and efficiency to users and enterprise customers. The spokesperson also said: Not all R&D projects will eventually be put into mass production, but this kind of all-round technology exploration is a key part of Google's full-stack AI layout. Why inference chips have become a hot spot in industry R&D A large number of companies are deploying AI inference chips (computing hardware that carries large models online), including startups such as SambaNova and d-Matrix, as well as OpenAI, And other technology giants. The goal of each company is to create inference chips that are more energy-efficient than NVIDIA GPUs and alleviate the industry pain points of global computing power shortage and high computing power costs. Nvidia itself also spent $20 billion in December last year to acquire related technology licenses from inference chip startup Groq. The technical routes of each manufacturer are different from Google's Frozen solution, but the design idea of Frozen v2 is similar to that of Canadian chip startup Taalas: both hard-code the computing logic of specific AI models into the chip hardware. Taalas has completed over US$200 million in financing from investors including Quiet Capital and Fidelity Investments. The design of Frozen v2 is derived from the original Frozen solution, which was led by Jeff Dean, chief scientist of Google DeepMind, and planned to directly burn the model weights into the chip hardware. However, this solution from many years ago had obvious shortcomings: the chip could only be adapted to a fixed version of the Gemini model, and the hardware service life cycle was too short, and it was eventually shelved.

Another person familiar with the matter revealed that the Google team is currently weighing the volume of solidified data built into the chip to find a balance between hardware energy efficiency and product versatility. Past earnings of Google’s self-developed chips Google’s years of self-developed chip layout has realized its value. This year, Google officially released the eighth-generation TPU and began selling the chip products directly to cloud customers, directly impacting Nvidia's monopoly in the AI chip market. Google has signed a multi-billion dollar TPU supply agreement with Meta, and is actively promoting its own TPU products to other cloud vendors that have been loyal to purchasing Nvidia hardware. Relying on self-developed chips and Gemini model co-design, Google's current hardware costs for running its own large models have dropped. However, the existing TPU only embeds a small amount of model underlying information, and its integration level is far less than that of the Frozen series chips. Google currently has no plans to mass-produce Frozen v2 chips, and its production capacity is far lower than TPU. The smaller mass production volume means that Google does not need to bring in external design and foundry partners early in the project. Google also regards this second generation of Frozen as a technical experiment: after the future AI large model architecture becomes finalized, the experience accumulated in this research and development can help engineers design more specialized high-performance AI chips.

← Back to archive