Crypto

Alphabet rises as Google reportedly designs a Gemini-specific AI chip

Google is reportedly working on Frozen v2, a custom server chip meant to run Gemini more efficiently as AI computing demand strains capacity.

Theo Nakamura

By Theo Nakamura · Staff Writer

· 3 min read

Alphabet rises as Google reportedly designs a Gemini-specific AI chip
Photo: Decrypt

Google is reportedly building a server chip tailored specifically for Gemini, a move that could cut the cost of running its flagship AI model. For investors, the key point is margin pressure: AI products can attract users fast, but every prompt also consumes expensive computing power.

The chip, code-named Frozen v2, was reported Monday by The Information, according to TechCrunch. Reports cited by Decrypt said Google engineers are aiming for a 2028 rollout and project the chip could be six to 10 times more efficient than Google’s current Tensor Processing Units.

Alphabet shares rose roughly 3% during Monday trading after the report, reaching $356 intraday, according to Decrypt. The move faded in the following session as investors waited for Alphabet’s Q2 2026 earnings, scheduled for Wednesday, July 22.

What Frozen v2 is meant to do

Google already uses custom AI chips called Tensor Processing Units, or TPUs. A TPU is a processor designed for machine-learning workloads, and Google has built them since 2015 to power Gemini and Google Cloud services used by outside developers.

Frozen v2 is reportedly different from a standard TPU upgrade. Decrypt reported that the chip would hardwire part of Gemini’s architecture into silicon. In AI, architecture means the design of the model, including how it routes and processes information. Putting part of that design directly into hardware can reduce repeated calculations and data movement between memory and processors.

The chip would not freeze Gemini’s “weights,” according to Decrypt. Weights are the learned values a model gains during training, which shape how it responds. Those would remain updatable, while the model’s structural blueprint would be fixed into the chip.

The efficiency claim centers on tokens per watt. Tokens are the small chunks of text an AI model reads and generates, while watts measure electricity use. If the reported engineering target holds, Google could produce more Gemini output for the same power bill.

Why capacity is part of the story

The reported chip effort comes as Google faces heavy demand for AI computing. Decrypt reported that in March, Google told Meta it could not supply the amount of Gemini compute Meta wanted to buy, and Meta told employees to ration AI usage.

Decrypt also reported that Google is spending up to $190 billion on AI infrastructure this year. Even with that spending, the company was reportedly turning away some demand because it lacked enough server capacity.

Cheaper AI inference, the process of running a trained model to answer user requests, could help Google serve more queries without increasing costs at the same pace. Decrypt reported that rival AI labs from China account for up to 45% of U.S. company AI token usage, in part because they operate 60% to 90% cheaper.

The Nvidia angle

The report also fits a broader push by large AI companies to rely less on Nvidia chips. Decrypt reported that Nvidia controls roughly 85% of the GPU market for AI. A GPU, or graphics processing unit, was originally built for rendering images and video games, but it has become central to training and running AI models.

Meta, Amazon, Microsoft and OpenAI all have custom silicon programs, according to Decrypt. The reason is straightforward: at the scale of large AI platforms, even small efficiency gains can change the cost of serving millions of prompts. A six-to-10-times improvement, if achieved, would be a significant shift in Google’s AI cost structure.

This story draws on original reporting from Decrypt.

More from Crypto

All Crypto