People familiar with the matter say Google’s newest chip promises a remarkable leap. Per unit of power, it could serve six to ten times more tokens than Google’s latest TPUs. Should everything proceed smoothly, Google aims to deploy the silicon as early as 2028. The Information first revealed the project, citing two people with direct knowledge.
Why pursue such a chip at all? The root cause lies in a punishing shortage of AI compute. Google has already struck a deal with SpaceX to rent computing capacity. That scarcity now hampers Google Cloud as well. Starved of capacity internally, Google has turned away some external customers. In effect, the company funnels nearly all available compute into its own products and services. Even so, the crunch persists. With no easy route to more capacity, investing in custom silicon becomes an attractive answer.
Etching Models Into Silicon Carries a Lifespan Problem
Conventional GPUs and TPUs work in a familiar way. First they load model weights into memory. Then the compute units read those weights back out to run inference. Consequently, the hardware must make time-consuming runtime decisions each time it meets a model.
Etching weights directly into silicon changes that dance entirely. Data simply flows along the circuits already carved into the wafer. Inference therefore accelerates dramatically.
Other companies have already tried baking models into silicon in pursuit of extraordinary throughput. Yet the approach carries an obvious flaw. Etched circuits cannot be altered afterwards. Once fabricated, the chip can never receive a model update. Models evolve at a furious pace today. As a result, hardwiring a model into silicon sharply shortens the chip’s useful life, retiring it long before its time.
Google’s Answer: Freeze the Architecture, Not the Weights
Google initially intended to etch the entire set of model weights into the chip. That was the Frozen v1 design, reportedly championed by DeepMind chief scientist Jeff Dean. Lifespan concerns, however, prompted Google to shelve it.
Frozen v2 takes a shrewder path. It freezes the architecture while leaving the weights updatable. Accordingly, the same chip can serve multiple generations of Gemini models.
Fewer Steps, Shorter Distances
Under Frozen v2, Google hardwires portions of Gemini’s architecture into the transistors themselves. This trims the number of steps the chip must execute. It also shortens how far data must travel per query. Throughput climbs substantially as a result. Sources suggest the design could slash inference latency enough to unlock entirely new applications.
One caveat deserves emphasis. The chip only stays useful while Google keeps building Gemini on the same underlying architecture. A shift away from today’s transformer paradigm would strand the hardwired portions. Moreover, Google has not yet settled on how much of the model to lock into silicon.
An Experiment, Not a Replacement
Google does not intend to manufacture Frozen v2 at TPU scale. Instead, it treats the chip as an experiment. The work lays groundwork for more specialised silicon once model designs settle. The project will therefore advance alongside Google’s existing TPU line rather than supplant it.
Google has offered a measured response to the reports. The company says its teams constantly explore and test new advances for performance and efficiency. Not every such effort, it notes, reaches production. Google has confirmed neither the specifications nor the timeline. Investors reacted warmly regardless, and Alphabet shares climbed after the news broke.
Support Our Threat Intelligence
If you find our technology report and cybersecurity news helpful, consider supporting our work.