Back to Home

Google's 'Frozen v2' Chip: 10x Gemini Efficiency by 2028

Google's parent company Alphabet is quietly developing a next-generation server chip internally codenamed "Frozen v2", designed specifically to supercharge the company's Gemini large language models. According to a report from The Information, the chip could deliver between six and ten times better efficiency than Google's existing AI accelerators when measured by tokens generated per unit of power. The chip is expected to reach production sometime in 2028.

This move places Google in an increasingly crowded field of AI companies designing their own custom silicon — a trend driven by the dual pressures of exploding compute demand and the urgent need to rein in costs. OpenAI unveiled its first custom inference chip, Jalapeño, in June, and Anthropic is reportedly in discussions with Samsung about a chipmaking partnership. But Google's approach is unique: it is co-designing hardware and software from the ground up, a strategy the company has been refining since its first Tensor Processing Unit (TPU) debuted in 2015.

What Is Frozen v2 and Why Does It Matter?

Frozen v2 is not a general-purpose GPU — it is a purpose-built accelerator optimized for transformer-based inference workloads. Google has not released architectural details, but based on the reported performance targets — 6–10x improvement in tokens per watt — we can infer several design choices. The chip likely employs a massively parallel systolic array architecture similar to Google's TPU lineage, but with significant enhancements tailored to the specific computational patterns of Gemini's mixture-of-experts (MoE) architecture.

The efficiency target is measured in tokens per watt, which is the most relevant metric for LLM inference at scale. To put 10x improvement in context: if Google's current TPU v5e delivers roughly 300–400 tokens per second at 250W per chip, a 10x efficiency gain would mean delivering the same throughput at 25W, or 10x the throughput at the same power envelope. Either scenario would dramatically reduce the cost of serving Gemini models in production.

Why Google Needs Its Own Silicon Now

Google's total planned capital expenditure of $180–190 billion in 2026 — disclosed earlier this year — underscores the urgency. At that burn rate, even modest efficiency gains translate into billions in operational savings. A 10x improvement in inference efficiency could slash the cost of serving Gemini by an order of magnitude, directly improving margins for Google's cloud business and enabling pricing that competitors cannot match.

Moreover, the chip effort is a hedge against NVIDIA's dominance. Google currently depends on NVIDIA GPUs for a significant portion of its AI compute, but the industry-wide push to reduce that dependence has accelerated. In June, reports indicated that even NVIDIA's largest customers are exploring alternatives. Google, with its deep experience in custom silicon from the TPU era, is arguably best positioned among the hyperscalers to break free.

The custom AI chip landscape reveals stark differences in strategy across the industry:

  • Google Frozen v2 (2028E): Inference-optimized transformer accelerator targeting 6–10x tokens/watt over TPU v5; full-stack co-design with Gemini models.
  • OpenAI Jalapeño (2026): Inference processor for GPT-series models; prioritizes latency over throughput; co-designed with Microsoft Azure.
  • Anthropic + Samsung (TBD): Partnership-based approach; leverages Samsung fabrication.
  • Amazon Trainium2 / Inferentia2 (shipping): AWS in-house chips for training and inference.
  • NVIDIA Blackwell Ultra (shipping): General-purpose AI GPU; dominant ecosystem lock-in.

Market Impact and Competitive Positioning

The market reacted immediately: Alphabet stock climbed ~3% on the news. This suggests investors see the chip program as a credible path to improving ROI on AI investments.

However, there are significant risks. Custom chip development is notoriously difficult — Google's Pixel lineup has cycled through multiple custom SoC efforts. A 2028 target also means the chip will arrive in a market that could look very different.

What is clear is that vertical integration in AI is accelerating. Google, OpenAI, Anthropic, Amazon, and Microsoft are all investing heavily in custom silicon. The winners of the next AI cycle will be determined not just by who builds the best models, but by who can serve them at the lowest cost.

The Bottom Line

Google's Frozen v2 chip represents a bet that the future of AI competition will be won on efficiency — not just intelligence. If the 6–10x tokens-per-watt targets are realized, the chip could reshape the economics of large-scale AI inference.

Comments

No comments yet. Be the first to share your thoughts!