Google is reportedly developing a custom inference chip codenamed Frozen v2, designed specifically for its Gemini AI model. According to a report by The New Stack, the chip would hardwire parts of Gemini’s architecture into silicon while keeping the model’s weights updatable. This compromise aims to deliver the efficiency of model-specific hardware without rendering the chip obsolete each time Gemini is updated. A Google spokesperson told The New Stack, “Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers.” Internal projections suggest the chip could achieve six to ten times more tokens per watt than Google’s current generation of AI chips.
Why This Matters: The ASIC-ification of AI Inference
The move signals a potential shift from general-purpose GPUs to model-specific ASICs for AI inference. Just as Bitcoin mining evolved from CPUs to GPUs to ASICs, AI inference may follow a similar path. Training still benefits from GPU flexibility, but once a model is in production, efficiency becomes paramount. Google’s Frozen v2 could relieve the AI compute crunch and significantly lower the cost of serving Gemini. For developers, this indicates that future AI systems will be designed with much tighter integration between model and hardware.
XPLAIN AI’s Analysis: Competitive Landscape and Market Impact
XPLAIN AI interprets this as part of a broader industry trend. Nvidia (NVDA) reportedly struck a $20 billion deal with Groq last year to license its inference technology. Canadian startup Taalas has demonstrated a chip embedding an entire 8-billion-parameter Llama model, claiming 17,000 tokens per second. d-Matrix offers an SRAM-based in-memory compute architecture, while SambaNova uses a custom dataflow approach. Google’s Frozen v2 occupies a unique middle ground, borrowing the hyper-efficiency of hardwired designs like Taalas while retaining enough flexibility for multiple product cycles.
Potential Beneficiaries and Risks
- Beneficiaries: Google (GOOGL) could see improved cloud margins and AI competitiveness. EDA tool vendors like Synopsys (SNPS) and Cadence (CDNS) may benefit from increased custom chip design activity.
- Risks: Nvidia (NVDA) faces potential erosion of its inference market share as ASICs gain traction. AMD (AMD) and Intel (INTC) are similarly exposed, though training workloads remain GPU-dominated for now.
Counter-Scenarios and Uncertainties
The project may not reach production. Google’s spokesperson noted that “not every project moves into production.” Rapid model updates could render a hardwired chip obsolete, and competitors are developing similar technologies. Key metrics to watch include the chip’s actual performance validation, production timelines, and the impact on Google’s cloud costs.
#Google #Frozenv2 #AIsemiconductors #ASIC #AIinference #Gemini #hardwired #chipdesign
Sources
- Google just bet its inference future on a chip built for one model — The New Stack | DevOps, Open Source, and Cloud Native News · News coverage · Mon, 20 Jul 2026 22:14:41 +0000
Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.