Gemini Architecture Written Directly Into Silicon: Technical Details of Google’s New AI Inference Chip “Frozen v2” Revealed, Processing per Unit of Power Consumption Increased by up to 10 Times

Source Tradingkey

TradingKey - According to The Information, citing two people familiar with the matter, Google ( GOOGL) is developing a brand-new server chip that can directly hardcode the underlying architecture of its Gemini large model into the chip's hardware, thereby significantly improving the operational efficiency of delivering AI services to users.

This AI inference chip, internally codenamed "Frozen v2", directly addresses Google's current pain point of a severe shortage of AI computing power. Sources revealed that the computing power gap has triggered internal resource conflicts, even forcing Google Cloud to turn down many external customer orders. The R&D team estimates that once the chip is officially deployed, the number of tokens processed per unit of power consumption will reach 6 to 10 times that of the latest generation of Google's existing self-developed AI chips.

1-2597621f464f428b8d557d5548b684e3

[Source: TradingView]

Stimulated by the news, Google's stock price performed strongly today, hitting an intraday high of $359.67. As of press time, it was still up 3.28% at $358.14, with the stock price returning above its 5-day, 10-day, and 20-day moving averages.

Why Is It Called Frozen?

The origin of the codename "Frozen" stems from its core design concept: permanently etching part of the Gemini large model's computing logic onto the silicon hardware—akin to "freezing" the software in the chip. Engineers are still finalizing the chip's core functional modules and the collaborative solutions for various components, and Google plans to deploy this AI chip as early as 2028.

The Frozen project aims to carve out a brand-new AI chip product line to complement Google's existing Tensor Processing Units (TPUs), rather than replace them.

Currently, the most widely used AI chips—Nvidia ( NVDA) GPUs and Google TPUs—are all general-purpose chips that can adapt to multiple different large models. However, when general-purpose chips run any model, they need to execute a large number of time-consuming logic judgments in real time; in contrast, the Frozen v2 AI inference chip directly builds in the underlying computing logic adapted to the Gemini series models, significantly reducing chip computing steps and data movement, thereby achieving a qualitative leap in inference efficiency.

People familiar with the matter said that this dedicated AI chip can respond to user queries faster, which is expected to support Google in launching entirely new AI application scenarios.

Google’s Strategic Bet on Frozen v2 Chip: Deep Integration of Hardware and Software

The technical roadmap of Frozen v2 means Google is making a strategic bet—wagering that it will continue to use the underlying architecture of the current Gemini large model. Only if subsequent iterations of the Gemini model adopt an underlying architecture consistent with the one used when the AI chip was designed can they run properly on the chip.

People familiar with the matter added that Google can still perform iterative optimization on the chip, including updating the built-in model weights—weights are the core parameter configurations that determine how the AI model answers questions.

A Google spokesperson responded that the company's teams "continuously research, develop, and test various innovative technologies, striving to achieve ultimate performance and energy efficiency for users and enterprise customers," adding that "not all R&D projects will eventually go into mass production, but this comprehensive technical exploration is a core part of our full-stack proprietary technology roadmap."

Google’s Software-Hardware Integration Ambition for Frozen v2 Chip

The revelation of the Frozen v2 AI chip reflects Google's deep strategic logic in the AI infrastructure sector, with AI compute anxiety being the primary driver. As the size of Gemini models continues to swell, the energy efficiency ceiling of general-purpose chips is drawing closer. Embedding large model logic directly into hardware represents a radical solution to break through the energy efficiency bottleneck of AI compute.

This also signals that Google is replicating an "Apple-style" vertically integrated path of combining software and hardware. From its self-developed TPUs to the Frozen series of dedicated AI chips, Google is attempting to establish a full-stack closed loop from model architecture to chip design in the AI era. If this path proves successful, it will drastically drive down AI inference costs, paving the way for Gemini's commercialization.

However, the risks are equally impossible to ignore. The deep coupling of chips and models means that once Gemini's underlying architecture undergoes a major iteration, Frozen AI chips could quickly become obsolete. Google's positioning of Frozen v2 as a "technical test platform" rather than a mass-production product precisely demonstrates that this path still requires time to be validated.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Will the Tech Rally Continue? The Technical Verdict on the NASDAQ 100 Riding a massive 32% post-earnings wave, the Nasdaq-100 is showing its first signs of exhaustion. We break down crucial exit and entry rules for long positions this week.
Author  Mitrade Team
6 Month 05 Day Fri
Riding a massive 32% post-earnings wave, the Nasdaq-100 is showing its first signs of exhaustion. We break down crucial exit and entry rules for long positions this week.
placeholder
Markets on a Wire: Imminent US Inflation Data Threatens to Lock In Fed Rate Hikes Imminent CPI and PPI data threaten to lock in a hawkish Federal Reserve rate hike cycle, leaving gold, tech equities, and Bitcoin highly vulnerable to a programmatic sell-off.
Author  Mitrade Team
6 Month 09 Day Tue
Imminent CPI and PPI data threaten to lock in a hawkish Federal Reserve rate hike cycle, leaving gold, tech equities, and Bitcoin highly vulnerable to a programmatic sell-off.
placeholder
XRP Price Prediction for July 2026: Can Buyers Finally Break the Downtrend?XRP (XRP) price trades near $1.05, caught between a year-long downtrend and a sudden burst of buying.July has historically rewarded XRP holders. This year the month arrives with on-chain accumulation
Author  Beincrypto
6 Month 30 Day Tue
XRP (XRP) price trades near $1.05, caught between a year-long downtrend and a sudden burst of buying.July has historically rewarded XRP holders. This year the month arrives with on-chain accumulation
placeholder
What to Expect From Ethereum (ETH) in July 2026Ethereum (ETH) enters July 2026 trading near $1,570, close to multi-month lows, after recording its first run of three consecutive red quarterly candles in its history.On-chain data and price charts n
Author  Beincrypto
7 Month 01 Day Wed
Ethereum (ETH) enters July 2026 trading near $1,570, close to multi-month lows, after recording its first run of three consecutive red quarterly candles in its history.On-chain data and price charts n
placeholder
NVIDIA Price Forecast: Michael Burry Shorts NVDA, but Analysts See $299On July 1, NVIDIA (NASDAQ: NVDA) sits at $198.34, failing to break above the former support level that is now serving as resistance between $198 and $205 on the 2H chart's downward blue c
Author  TradingKey
7 Month 02 Day Thu
On July 1, NVIDIA (NASDAQ: NVDA) sits at $198.34, failing to break above the former support level that is now serving as resistance between $198 and $205 on the 2H chart's downward blue c
goTop
quote