Google develops new “Frozen” dedicated chip to significantly improve AI model operating efficiency
According to two people familiar with the matter, A new server chip is being developed, which can directly solidify the underlying architecture of the Gemini large model into the chip hardware, thereby greatly improving the operating efficiency of providing AI services to users. This chip, internally codenamed Frozen v2, is designed to solve Google's current serious shortage of AI computing power. The computing power gap has caused internal resource conflicts and even forced Google Cloud to reject cooperation orders from many external customers. People familiar with the matter said that the R&D team calculated that after the chip is officially launched, the number of tokens that can be processed per unit power consumption will reach 6 to 10 times that of the latest generation of Google's existing self-developed AI chips, achieving a qualitative leap in energy efficiency. The origin of the code name "Frozen" comes from its core design idea: permanently etching part of the computing logic of the large model into the silicon hardware. Engineers are still finalizing the core functional modules of the chip and how each component will work together. People familiar with the matter said that Google plans to deploy the chip as early as 2028. Project Frozen aims to create a new self-developed chip product line that complements Google's existing tensor processors (TPUs), rather than replaces them. The most widely used AI chips at present, NVIDIA GPU and Google TPU, are general-purpose chips that can be adapted to a variety of different large models. However, when general-purpose chips run any model, they need to perform a large number of time-consuming logical judgments in real time; Frozen v2 directly has built-in underlying computing logic that adapts to the Gemini series models, which can significantly reduce chip computing steps and data transfer volume. One of the people familiar with the matter said that this special chip can respond to user questions faster and is expected to support Google's launch of new AI application scenarios. But this chip also means that Google is betting that it will continue to use the underlying architecture of the current Gemini large model: only subsequent iterations of the Gemini model adopt the same infrastructure as the chip was designed to run normally on the chip. People familiar with the matter added that Google can still iteratively optimize the chip, including updating the built-in model weights (weights are the core parameter configuration that determines how the model answers questions). Google's official spokesperson responded: The company's teams "continue to develop and test various innovative technologies, striving to achieve ultimate performance and energy efficiency for users and enterprise customers." The spokesperson added: "Not all R&D projects will be put into mass production, but this kind of all-round technology exploration is a core part of our full-stack self-research technology route." Why inference chips have become the focus of the industry Including start-ups such as Samba Nova and d-Matrix, as well as OpenAI, Technology giants such as AI are developing AI inference chips - computing hardware specifically used to provide external services for large models. The goal of each company is to achieve inference energy efficiency that exceeds that of NVIDIA GPUs, alleviate the shortage of computing power, and reduce operating costs. Nvidia itself also spent US$20 billion in December last year to acquire the technology license of inference chip manufacturer Groq. The technical routes of various players in the industry are different from Google's Frozen solution, but the design idea of Frozen v2 is highly similar to that of Canadian chip startup Taalas: the latter also chose to hard-code the logic of specific large models into the chip. Taalas has completed over US$200 million in financing, with investors including Quiet Capital and Fidelity Investments. The development iteration of Frozen v2 comes from the original Frozen solution. The first-generation solution is led by Jeff Dean, chief scientist of Google DeepMind, and plans to etch the complete model weights directly into the chip. However, this solution has obvious shortcomings: the chip can only adapt to a single version of Gemini, and the hardware life cycle is extremely short, so it has been shelved for many years. Another person familiar with the matter revealed that Google is still weighing how much model information is solidified in the Frozen v2 chip to balance hardware flexibility and operating energy efficiency. Google’s deployment of self-developed chips has reaped rewards so far. This year, Google launched the eighth-generation TPU and promoted it directly to cloud customers, directly impacting Nvidia's monopoly market share. Google has signed a multi-billion dollar deal with Meta to lease TPU computing power to it, while trying to supply TPU to Nvidia’s core cloud vendor customers.
Self-developed AI chips and not relying entirely on Nvidia GPUs have helped Google reduce the operating costs of the Gemini model - Google adapted it around Gemini when designing the TPU, but the model information built into the TPU is far less than that of the Frozen series chips. Google has no plans to mass-produce Frozen v2 at the moment, and the production capacity will be much lower than TPU. People familiar with the matter said that the mass production scale of this chip is limited, and Google can postpone the introduction of external design and foundry partners; at the same time, Google regards this generation of products as a technology test platform. After the underlying architecture of the large model stabilizes, it will accumulate engineering experience for the development of more specialized chips.
Self-developed AI chips and not relying entirely on Nvidia GPUs have helped Google reduce the operating costs of the Gemini model - Google adapted it around Gemini when designing the TPU, but the model information built into the TPU is far less than that of the Frozen series chips. Google has no plans to mass-produce Frozen v2 at the moment, and the production capacity will be much lower than TPU. People familiar with the matter said that the mass production scale of this chip is limited, and Google can postpone the introduction of external design and foundry partners; at the same time, Google regards this generation of products as a technology test platform. After the underlying architecture of the large model stabilizes, it will accumulate engineering experience for the development of more specialized chips.