Google is developing a new server chip for Gemini that could make its in-house AI models far more efficient, a move that underscores how urgently the company and its rivals are chasing lower-cost inference at scale. The chip, reportedly called Frozen v2, is expected to arrive around 2028 and could significantly reduce the power needed to generate responses from Google’s models.
The effort matters because AI is becoming an infrastructure business as much as a software one. As companies race to deploy larger models to more users, the cost of running those systems has become a central competitive issue, shaping investment plans, profit margins, and the balance of power between model builders and chip suppliers.
What Google’s Frozen v2 chip is meant to do
Frozen v2 is being designed to help Gemini run more efficiently inside Google’s own data centers, according to a report from The Information. The chip is aimed at improving the economics of inference, the stage where a model processes prompts and generates answers for users.
That distinction matters. Training a model gets most of the early attention, but inference is what drives ongoing operating costs once a chatbot or AI assistant is in the wild. If Google can reduce the energy and compute required for each token generated, it can lower expenses across Gemini-powered products while creating more room to scale usage.
How much more efficient could it be?
According to the report, Google believes the chip could be six to 10 times more efficient than its current AI hardware, measured by tokens generated per unit of power. That is a striking claim, though it remains an early-stage projection rather than a public product specification.
If that performance target proves accurate, it would mean Google could squeeze much more work out of the same power budget. In practical terms, that would help the company serve more queries, reduce electricity demand, and improve the unit economics of Gemini deployments.
Google said its teams are continually testing new ideas to improve performance and efficiency, while noting that not every experiment reaches production. The company described its hardware-and-software approach as a way to keep systems tightly integrated and optimized for real-world workloads.
Why AI companies are racing to build custom chips
Google is hardly alone in trying to design silicon tailored to its own AI stack. Across the industry, major AI developers are looking for ways to reduce dependence on outside chip suppliers and gain tighter control over cost, performance, and supply.
The most obvious reason is economics. General-purpose accelerators can be powerful, but they are not always the most efficient fit for a company’s own models and workloads. Custom chips can be tuned for specific tasks, which can translate into lower latency, less power draw, and better throughput.
There is also a strategic reason. Nvidia has dominated the AI accelerator market, and that dominance has left many of the biggest AI players dependent on its hardware and supply cadence. By building more of their own chips, companies hope to create leverage, reduce bottlenecks, and bring more of the value chain in-house.
The AI chip race is intensifying
The competition is no longer limited to model quality or app features. It now includes who can deliver the same intelligence at the lowest possible cost per query.
Recent moves across the sector show how quickly this race is escalating:
- OpenAI unveiled its first custom inference chip, called Jalapeño, in June.
- Anthropic was reported earlier this month to be discussing a chipmaking partnership with Samsung.
- Google is already pursuing multiple generations of proprietary hardware to support its AI services.
That pattern suggests the market is moving toward vertically integrated AI stacks, where the biggest players control both the model software and the silicon running it.
Why the chip matters for Alphabet’s spending plans
Alphabet’s massive capital commitments have become a focal point for investors, especially as the company pours money into AI infrastructure, data centers, and custom hardware. Earlier this year, Google said it expected to spend between $180 billion and $190 billion, a figure that has raised questions about how quickly those investments can translate into revenue and competitive advantage.
In that context, a more efficient chip is not just a technical milestone. It is also a financial message. If Google can show that its AI spending is producing better unit economics, it may be able to ease concerns that the company is overbuilding for demand that may take time to fully materialize.
Investors have been looking for proof that the heavy spending spree will pay off. A custom chip that cuts inference costs would support the argument that Google can build an AI business with stronger margins than a model reliant on rented or externally sourced compute.
| Company | Custom AI chip move | Likely focus | Why it matters |
|---|---|---|---|
| Frozen v2 | Gemini inference efficiency | Could lower power use and operating costs in Google’s own AI stack | |
| OpenAI | Jalapeño | Inference processing | Signals a push to reduce dependence on third-party hardware |
| Anthropic | Reported Samsung talks | Chipmaking partnership | Suggests the industry is broadening beyond model development |
How does this fit Google’s long-term AI strategy?
Google has been one of the earliest and most aggressive companies in combining software, models, and infrastructure under one roof. The company’s approach has long centered on designing the stack end to end, from cloud systems to model deployment to specialized chips.
That integrated model gives Google a potential advantage if it can execute well. Instead of relying solely on off-the-shelf hardware, it can tailor its chips to the specific needs of Gemini, Search, cloud customers, and other internal workloads.
By doing so, Google can aim for performance gains that are difficult for rivals to copy quickly. It can also potentially improve reliability and reduce exposure to supply shortages or pricing swings in the broader chip market.
What is inference, and why is it so important?
Inference is the phase where an AI model takes an input and produces a result, such as a chatbot reply, image description, code suggestion, or search summary. It is the part of AI that users actually experience, and it is also the part that can become expensive at scale.
That is why chip design has become such a strategic issue. A model that is only slightly better but dramatically cheaper to run can be more commercially attractive than a more powerful one with high operating costs. In a crowded market, efficiency can become a decisive advantage.
What the stock market is reading into the report
Shares of Alphabet rose about 3% on Monday morning after The Information’s report, suggesting investors viewed the news as a positive sign ahead of the company’s upcoming earnings report. The market reaction indicates that Wall Street is not only watching product launches, but also scrutinizing how well Google can convert AI investment into financial discipline.
The optimism is understandable. If a future chip can materially reduce the cost of running Gemini, that could help Alphabet defend margins while expanding AI usage. It also offers a counterpoint to the narrative that AI spending is simply an endless drain on cash.
Still, the reaction should be treated cautiously. A reported chip project is not the same as a shipping product, and efficiency gains at the design stage do not always survive manufacturing, deployment, and scaling.
How Google’s chip push compares with the broader market
Google’s move fits a larger industry shift toward custom silicon, but the company arrives with a unique advantage: it already has extensive experience designing AI-focused chips, including the Tensor Processing Unit family used across its infrastructure.
That background gives Google a head start over newer entrants. It also means the company has a clearer feedback loop between model behavior and hardware design, which is essential when the objective is not just raw performance but performance per watt.
At the same time, the business case is becoming more urgent. As AI services become mainstream, the winners may be the companies that can serve the most users for the least amount of compute. In that environment, every gain in efficiency compounds.
Potential implications for Gemini
If Frozen v2 reaches production and meets expectations, Gemini could become cheaper to operate at scale. That would help Google in several ways:
- Lower operating costs for consumer AI products.
- Better economics for enterprise and cloud AI offerings.
- More headroom to expand usage without proportionally increasing power consumption.
- Greater independence from external chip supply and pricing.
Those benefits would not arrive overnight, and a 2028 timeline leaves plenty of room for the market to shift. But the direction is clear: Google wants the chips powering its AI systems to be as strategically important as the models themselves.
Timeline of key developments
| Date | Event | Why it matters |
|---|---|---|
| Earlier this year | Google said it planned to spend $180 billion to $190 billion | Reinforced investor focus on whether AI spending will pay off |
| June 2026 | OpenAI unveiled its first custom chip, Jalapeño | Showed the chip race is widening beyond Google and Nvidia |
| Earlier this month | Anthropic was reported to be discussing Samsung chipmaking talks | Suggested more AI firms are seeking hardware partnerships |
| July 20, 2026 | Report surfaced on Google’s Frozen v2 chip | Prompted a stock gain and renewed interest in Google’s AI strategy |
| 2028, expected | Possible release window for Frozen v2 | Indicates a long-term bet on custom AI infrastructure |
What happens next?
For now, Google is not confirming a product launch, and the chip remains a reported project rather than an official release. The company’s public comments suggest experimentation is ongoing, which is typical for major semiconductor efforts that can take years to mature.
Over the coming quarters, investors will likely watch for clues in Alphabet’s earnings calls, cloud disclosures, and AI product updates. Any sign that Google is improving Gemini’s efficiency at the hardware level could strengthen confidence in its long-term AI economics.
More broadly, the report reinforces a simple reality: the next phase of AI competition will be fought not only in model benchmarks, but in watts, tokens, and data-center economics. Google is now signaling that it intends to compete on all three.
Background: why this moment matters for the AI industry
The chip announcement arrives at a time when the AI industry is moving from hype toward operational discipline. The first wave of excitement centered on model capability. The next wave is about whether those capabilities can be delivered profitably, reliably, and at global scale.
That shift favors companies with deep infrastructure resources. Google, Microsoft, Amazon, Meta, and others have the balance sheets to invest heavily in hardware and data centers. But even among those giants, the firms that can extract the most useful work from each watt of electricity may have the edge.
In that sense, Frozen v2 is more than a single chip rumor. It is a sign that Google sees AI efficiency as one of the decisive battlegrounds of the next several years.
If the chip performs as described, it could help shape how Google prices Gemini, how aggressively it expands AI features, and how confidently investors view Alphabet’s enormous spending program. If it falls short, it will still reflect the direction of travel across the industry: AI is becoming a hardware story as much as a software one.









