Google Builds “Frozen v2,” an In-House Chip to Run Gemini Far More Efficiently


Google is reportedly working on a new server chip, internally called “Frozen v2,” designed to run its Gemini AI models with unprecedented efficiency. The idea, revealed on July 20, 2026 by specialist outlet *The Information* and then picked up widely, is as bold as it is radical: etch Gemini’s very architecture directly into silicon. The hoped-for result, according to the project’s engineers: 6 to 10 times more tokens generated per watt than Google’s current chips. It’s a direct answer to the compute capacity crunch now straining the whole AI sector — and a strong signal in Google’s battle with Nvidia.

puce-google-v2 Google Builds "Frozen v2," an In-House Chip to Run Gemini Far More Efficiently

The principle: “freezing” the model into the chip

Most AI chips process a model stored separately in memory, making a host of dynamic decisions with every response. Frozen v2 would take another path: embedding Gemini’s architecture blueprint directly into the hardware. By “freezing” this structure into the silicon, the chip cuts the number of calculations to perform and shortens how far data travels, two major sources of energy waste when generating a response.

An important nuance, and the whole point of the approach: it’s the broad architecture that would be etched, not the “weights” (the tuned parameters that encode what the model has learned). Google would thus keep some flexibility to evolve Gemini while gaining enormous efficiency — a trade-off observers call “flexible hardwiring.”

6 to 10 times more tokens per watt

The figure floated by engineers is spectacular: 6 to 10 times more tokens generated per unit of energy than Google’s latest in-house chips. Since the “token” is the basic unit an AI model processes (a fragment of a word), this “tokens per watt” metric directly reflects the energy efficiency of inference — the phase where the model answers your queries. In plain terms: at equal power draw, a Frozen v2 chip could serve far more users.

A complement to TPUs, not a replacement

Frozen v2 wouldn’t spell the end of Google’s famous TPUs (Tensor Processing Units), its in-house AI chips. It would rather be a specialized branch of its portfolio, entirely dedicated to Gemini inference, while TPUs remain general-purpose processors that can also train models. For reference, Google recently deployed Ironwood, its seventh-generation TPU built for “the age of inference.” Frozen v2 would add to this arsenal, with deployment targeted for 2028 — engineers still finalizing the design and exactly how much information to freeze into the silicon.

The compute crunch, driving the project

The context is decisive. Google faces an internal compute capacity crunch so severe it reportedly created tensions within the company and pushed Google Cloud to decline external customers. The pressure is such that Google reportedly agreed, in June 2026, to pay SpaceX nearly $920 million a month for access to roughly 110,000 Nvidia GPUs as bridge capacity for its Gemini Enterprise platform. Designing an ultra-efficient chip tailored to Gemini aims precisely to ease that squeeze and cut an energy bill that has become colossal.

Challenging Nvidia and controlling costs

Beyond efficiency, Frozen v2 is part of a broader strategy: reducing dependence on Nvidia, whose GPUs dominate the AI market, and lowering the operating cost of Gemini. By controlling its silicon end to end, Google hopes to serve its AI cheaper and at greater scale, while competing with Nvidia on cloud services. On the report, shares of Alphabet, Google’s parent, rose — a sign the markets see this efficiency as a major competitive edge.

Why 2028, and why stay cautious

A 6-to-10x efficiency factor is impressive, but keep a cool head: these are still internal estimates for a chip expected in 2028. The design isn’t frozen, and the exact amount of information to etch into the silicon remains to be settled. As often with hardware announcements, the gap between the promise and industrial reality can be significant. We’ll follow this story closely.

The giants’ race for in-house silicon

Google isn’t alone on this front. For two years now, every tech giant has been designing its own AI chips to reduce dependence on Nvidia and control costs. Amazon pushes its Trainium and Inferentia chips for training and inference; Microsoft develops its Maia family; Meta advances with its MTIA accelerators; and Apple has always designed its own processors for its devices. In this landscape, Frozen v2 marks a radicalization of the approach: where the others design generic AI chips, Google would etch an architecture specific to its flagship model.

This extreme specialization is a gamble. It promises unmatched efficiency gains, but at the cost of rigidity: a chip tailored for Gemini serves another model less well. It’s a coherent choice for a player who, like Google, controls the model, the software and the hardware — a vertical-integration advantage few rivals can claim.

Weights vs architecture: why the distinction matters

To grasp what makes Frozen v2 original, it helps to separate two things a model is made of. The architecture is the blueprint — how the layers are organized, how information flows through them. The weights are the millions or billions of tuned values the model learned during training; they’re what makes one model “smarter” than another and what gets updated with each new version.

Etching the architecture into silicon while keeping the weights flexible is the clever compromise. It bakes in the stable, structural part (which changes rarely) and leaves adjustable the part that evolves often. That’s why Google can hope for huge efficiency gains without locking itself into a single, frozen version of Gemini — a balance that pure “hardwired” chips of the past never managed.

What it means for you

You’ll never buy a Frozen v2 chip: it will stay in Google’s data centers. But its impact could be very real for users. An AI that’s less energy-hungry potentially means faster, cheaper and more available services — and a better-controlled environmental footprint, as data centers’ electricity and water use worries a growing number of people. It’s also a strong signal in the AI chip war, where Google, Nvidia, Amazon and others compete to stop depending on a single supplier. And for everyday users, cheaper inference could eventually mean AI features that today feel premium becoming free, or usage limits on your favorite assistant loosening over time.

The environmental stake deserves a pause. The explosion of generative AI has sharply increased data centers’ electricity and water use, to the point of becoming a political issue in several regions. A chip able to generate 6 to 10 times more responses for the same energy wouldn’t solve everything — overall usage keeps growing — but it would ease pressure on the grid and on each query’s carbon bill. For Google, whose climate targets have been strained by the AI race, Frozen v2’s efficiency is as much an economic imperative as an environmental argument.

To go further, find our other analyses on artificial intelligence and hardware: our AI and hardware news on Wanda-techs.

Does this in-house AI chip race look decisive to you? Tell us in the comments on Wanda-techs.com.

Share this content:

Ingénieur passionné et rédacteur web depuis 2018, j'allie mon expertise technique à ma passion pour l'écriture pour partager astuces, actualités et savoirs pratiques avec la communauté.

Post Comment

Vous avez certainement manqué...