Tue 21 Jul 2026 / 11:46 ET
Kernel
Hardware 3 min read

Google is said to be developing a Gemini-specific AI server chip

The Information says Google's Frozen v2 project could hardwire parts of Gemini into silicon to improve inference efficiency by 2028.

Felix Aranda

By Felix Aranda / Silicon Editor

Google is said to be developing a Gemini-specific AI server chip
img: Tom's Hardware

Google is working on a server processor designed around Gemini, according to The Information, which cited two people with direct knowledge of the effort. The project, informally called Frozen v2, is aimed at making AI inference less power-hungry by putting some of Gemini’s model design into the chip itself.

The claimed upside is large, although still just a projection. Engineers on the project expect Frozen v2 could generate six to 10 times more tokens per watt than Google’s newest tensor processing units, The Information reported. Deployment could come as soon as 2028, according to the report.

That target matters because inference is where AI systems spend money every time users ask for an answer. The two people cited by The Information said the work is partly driven by a compute crunch severe enough that Google Cloud has declined some deals with outside customers. Google has not publicly confirmed the chip.

The technical bet is specialization. A TPU, like a GPU, is general enough to run many models loaded onto it. That flexibility has a cost: the chip has to make runtime choices and move data around as it executes the model. Frozen v2 would fix some Gemini-specific behavior in hardware, according to the report, cutting down the work the chip has to do for each query.

One person familiar with the project told The Information that the design could lower latency enough to support new kinds of applications. That is a claim about a chip that does not yet appear to exist as a deployed product, so treat it as an engineering goal rather than a benchmark.

The project is a scaled-back version of an earlier Frozen design led by Google DeepMind chief scientist Jeff Dean, The Information reported. That first proposal would have embedded Gemini’s weights, the model parameters learned during training, into silicon. Google shelved that approach because a chip tied to one model version would age badly.

Frozen v2 is meant to hardwire architecture instead of weights. That would let Google update Gemini’s weights while keeping the chip usable across model releases, as long as later Gemini models keep the same basic architecture. The Information said Google has not yet decided how much of the model design to lock into the hardware.

The chip is not expected to replace Google’s TPU line. The Information reported that Google does not plan to build Frozen v2 at TPU-scale volumes and sees this generation partly as a test for more specialized chips if model designs become more stable. Google’s eighth-generation TPUs, announced at Cloud Next in April, split the line into separate training and inference variants.

The reported 2028 timing also lines up with Google’s broader accelerator buildout. The company has reportedly booked Intel to package more than 3 million TPUs that year.

Other companies are also chasing model-specific inference hardware. Taalas, a Toronto startup backed by investors including Quiet Capital and Fidelity, launched its HC1 chip in February with Llama 3.1 8B wired into an 815-square-millimeter die on TSMC’s N6 process. Taalas claims the chip can deliver 17,000 tokens per second per user without HBM on the package. Nvidia struck a $20 billion deal in December to license technology from Groq, another inference-chip designer.

A Google spokesperson told The Information that not every project becomes a product and that “this rigorous exploration is central to our full stack approach.”

This story draws on original reporting from Tom's Hardware.

More Hardware/

view all ↗