Skip to main content
AI HardwareEfficiency Leap· 4 min read· in Artificial Intelligence

Google's 'Frozen v2' Custom AI Chip Promises 10x Efficiency Gain for AI Inference

Alphabet is reportedly developing a specialized server chip that hardwires its Gemini AI architecture directly into silicon, aiming to drastically cut the energy costs of running generative AI.

By Karim Mansour

Hardware Engineers 30%Cloud Infrastructure Providers 30%Financial Analysts 25%Industry Competitors 15%
Hardware Engineers
Focusing on the technical milestone of embedding software architecture into physical silicon.
Cloud Infrastructure Providers
Focusing on solving the compute capacity crunch and reducing the massive energy footprint of AI.
Financial Analysts
Evaluating the chip as a necessary mechanism to justify Alphabet's massive capital expenditures.
Industry Competitors
Focusing on the broader race to build custom silicon and bypass general-purpose GPU suppliers.

Perspectives this story doesn't cover

  • Environmental Advocates monitoring the absolute energy consumption of AI data centers.
  • Enterprise Cloud Customers waiting for compute availability.

Running state-of-the-art artificial intelligence models requires an astounding amount of electricity, and processing every user prompt—a step known as inference—quickly adds up to massive operational costs. Now, Google is reportedly on the verge of a hardware breakthrough that could fundamentally rewrite the economics of generative AI.[1][2]

Alphabet, Google's parent company, is developing a highly specialized server chip internally codenamed "Frozen v2," designed specifically to run its in-house Gemini AI models with unprecedented efficiency.[2][3]

According to reports from The Information and TechCrunch, the new silicon aims to deliver a staggering six to ten times improvement in power efficiency compared to Google's current generation of custom AI accelerators.[1][2]

This efficiency is measured by a crucial industry metric: tokens generated per unit of power, or "tokens per watt." A tenfold leap in this metric would allow Google to serve complex AI queries at a fraction of today's energy footprint, fundamentally altering the unit economics of the technology.[3][5]

The Frozen v2 chip targets a massive leap in tokens generated per unit of power.

The engineering approach behind Frozen v2 is genuinely unusual. Rather than running Gemini's model weights and computations on general-purpose AI accelerator silicon, Google plans to hardcode specific elements of the model's architecture directly into the chip's physical design.[4][7]

Traditional accelerators, including Google's own Tensor Processing Units (TPUs), remain programmable to run any model architecture. They are flexible calculators. In contrast, the Frozen v2 chip locks specific neural network logic directly into the silicon logic.[4][7]

This direct hardware integration drastically cuts down computational steps. It minimizes the constant data movement between processors and memory modules during inference, which is typically where the most power is lost. By turning software design into physical circuits, the system achieves massive throughput without drawing excess power.[4][7]

By embedding software logic into physical circuits, Google aims to eliminate data-movement bottlenecks.
This direct hardware integration drastically cuts down computational steps.

The project represents a distinct, parallel branch from Google's existing TPU lineup. The company's latest generation consists of the TPU v8t for training and TPU v8i for inference. Frozen v2 isn't meant to replace those chips; instead, it serves as a specialized line focused entirely on efficient model serving in fixed production environments.[2][4]

This extreme specialization comes at a critical time for Google Cloud. The sheer demand for AI computing capacity has fueled internal resource tensions and reportedly prompted Google Cloud to decline deals with outside enterprise customers due to a lack of available compute.[1][6]

Developing ultra-efficient internal hardware directly addresses that bottleneck. By swapping out general-purpose chips for highly optimized silicon, Google can process significantly more workloads within its existing data centers without overloading local electric grids or requiring massive new facility build-outs.[3][6]

The financial stakes for Alphabet are enormous. The company has guided that it plans to spend between $180 billion and $190 billion on capital expenditures in 2026 alone, largely driven by AI infrastructure build-outs.[3][5]

Alphabet's massive capital expenditures are driving the need for extreme hardware efficiency.

This eye-popping figure has raised eyebrows among investors concerned about the long-term profitability of generative AI. News of the Frozen v2 chip appears to have assuaged some of those worries, demonstrating a clear path to lowering unit costs. Following the initial reports, Alphabet's stock climbed roughly 3%.[3][5]

However, the "flexible hardwired" strategy carries significant engineering risks. Deep software-hardware coupling means that if the underlying architecture of the Gemini model undergoes a major paradigm shift before the chip's release, the silicon risks becoming obsolete before it even hits the server racks.[7]

To mitigate this, Google has reportedly shifted from an earlier "fully hardwired" approach that baked in specific model weights to a strategy that only hardwires the broader architecture, extending the chip's usable lifecycle while still capturing massive efficiency gains.[7]

Extreme hardware efficiency allows data centers to process more workloads without overloading local power grids.

The Frozen v2 project reflects a wider shift seen across the tech industry. As the size of frontier models continues to swell, the energy efficiency ceiling of general-purpose chips is drawing closer. Companies are increasingly building their own specialized silicon to reduce dependence on Nvidia, which has dominated the AI hardware supply chain for years.[1][7]

OpenAI recently released its first custom inference processor, dubbed "Jalapeño," designed in collaboration with Broadcom. Meanwhile, Anthropic is reportedly discussing a new chipmaking partnership with Samsung to optimize its Claude models.[1][5]

Google's Frozen v2 is currently slated for deployment as early as 2028. While engineers are still finalizing the design and determining the precise amount of model information that will be hardwired, the project signals that the next era of AI competition will be won not just in software labs, but in the physical architecture of the silicon itself.[2][6]

Key takeaways

  • Google is developing a custom server chip, codenamed 'Frozen v2,' designed specifically to run its Gemini AI models.
  • The chip hardwires parts of the AI model's architecture directly into the silicon, rather than running it as software.
  • Engineers project the chip will deliver a 6 to 10 times improvement in inference efficiency, measured in tokens per watt.
  • The extreme efficiency aims to solve severe compute capacity shortages that have forced Google Cloud to turn away customers.
  • The project highlights a broader industry shift toward highly specialized inference chips to reduce reliance on general-purpose GPUs.

Terms in play

Inference
The process of a trained AI model generating responses or predictions based on new data, such as answering a user's prompt.
Tokens per watt
A metric measuring how much AI output is generated for every unit of electrical power consumed.
Hardwiring
The practice of embedding software instructions or model architecture directly into the physical circuitry of a microchip.
Tensor Processing Unit (TPU)
Google's custom-designed application-specific integrated circuit (ASIC) used to accelerate machine learning workloads.

Sources

Source coverage

7 outlets

4 viewpoints surfaced

Hardware Engineers 30%Cloud Infrastructure Providers 30%Financial Analysts 25%Industry Competitors 15%
  1. [1]TechCrunchCloud Infrastructure Providers

    Google reportedly developing new server chip to run Gemini more efficiently

    Read on TechCrunch
  2. [2]The InformationIndustry Competitors

    Google Designs Server Chip to Lower Cost of Running AI Models

    Read on The Information
  3. [3]Android HeadlinesFinancial Analysts

    Google's Next AI Chip Aims to Supercharge Gemini with Up to 10x Efficiency

    Read on Android Headlines
  4. [4]Crypto BriefingHardware Engineers

    The 'Frozen v2' chip promises 6-10x efficiency gains over current TPUs

    Read on Crypto Briefing
  5. [5]Tech in AsiaFinancial Analysts

    Alphabet designing server chip for Gemini

    Read on Tech in Asia
  6. [6]India TimesCloud Infrastructure Providers

    Google Plans New AI Chip for Gemini Models, Targets Efficiency Gains

    Read on India Times
  7. [7]BigGoHardware Engineers

    Google Bets on 'Flexible Hardwired' Chips: Gemini Architecture Etched into Silicon

    Read on BigGo

Comments

Stay informed

Every angle. Every day.

Get Artificial Intelligence stories with full source coverage and perspective breakdowns delivered to your inbox.