Skip to content
KO EN
AI 기술 Upcoming

Google Deprecates Gemini 3.5 Flash After Just Two Months, Unveils 3.6 Flash as 3.5 Pro Remains Delayed

Google has made a surprising move in its AI model lineup, announcing the deprecation of Gemini 3.5 Flash —the model it touted as the star of its I/O event

Google has made a surprising move in its AI model lineup, announcing the deprecation of Gemini 3.5 Flash—the model it touted as the star of its I/O event in May—after only two months on the market. In its place, the company has released Gemini 3.6 Flash, a faster and cheaper version that Google says incorporates user feedback from the 3.5 release. Alongside this, Google introduced two other new models, including its first cybersecurity-specific Gemini variant. However, the much-anticipated Gemini 3.5 Pro, originally slated for a June launch, remains in testing with no new release date announced. This rapid generational shift, combined with the delay of the flagship model, sends mixed signals about Google’s AI strategy.

What Happened: A Swift Succession and a Missing Pro Model

According to Google, the changes in Gemini 3.6 Flash are driven by user feedback, particularly around code generation quality, where the 3.5 Flash reportedly fell short of expectations. The new model is described as marginally more capable with improved multimodal features, but the emphasis is on efficiency: Google claims it uses up to 65 percent fewer tokens than its predecessor, addressing growing enterprise concerns over AI token costs. The simultaneous launch of a cybersecurity-focused Gemini model signals Google’s push into vertical-specific AI solutions. Yet the absence of Gemini 3.5 Pro, which was expected to compete with frontier models from OpenAI and Meta, raises questions about Google’s ability to deliver top-tier performance in a timely manner. It is important to note that Google has not officially explained the delay, and the company may simply be adjusting its release schedule.

Why It Matters: Cost Efficiency Becomes the New Battleground

Google’s pivot to a cheaper, faster model highlights a broader shift in the AI industry: the focus is moving from raw performance to practical, cost-effective deployment. As businesses become increasingly sensitive to the expense of AI tokens, Google is positioning Gemini 3.6 Flash as a solution for high-volume, budget-conscious applications. This strategy could help Google capture market share among small and medium-sized enterprises that are hesitant to adopt expensive AI services. Meanwhile, the delay of Gemini 3.5 Pro may allow competitors like OpenAI and Microsoft to strengthen their positions in the premium AI segment. However, the cybersecurity model opens a new front, potentially creating a specialized niche for Google in security analytics and threat detection.

XPLAIN AI’s Interpretation: Signals from the Pro Delay and Market Implications

XPLAIN AI interprets the delay of Gemini 3.5 Pro as more than a scheduling hiccup. It may indicate that Google is reallocating resources toward efficiency and specialized models rather than chasing the frontier performance race. In the short term, this could benefit Microsoft’s Azure and OpenAI’s GPT-4o, which continue to lead in high-end capabilities. However, Google’s long-term bet on cost efficiency could pay off if the market shifts toward value-driven AI adoption. The cybersecurity model also suggests Google is diversifying its AI portfolio to reduce reliance on general-purpose models. Still, the uncertainty around Gemini 3.5 Pro‘s release means Google’s competitiveness in the frontier model space remains an open question. Investors should watch for independent benchmarks of Gemini 3.6 Flash and any updates on the Pro model’s timeline.

Potential Beneficiaries and Risks in the Ecosystem

Based on the technology’s implications, the following stakeholders could be affected:

  • Potential beneficiaries: Google (GOOGL) itself may benefit as its cloud customers gain access to cheaper AI inference, potentially driving adoption. AI startups focused on cost-sensitive applications could also gain from lower token prices. Additionally, the cybersecurity model may create new opportunities for Google’s security partners.
  • Potential risks: Companies reliant on high-performance AI models, such as those competing with OpenAI, may face pressure if Google’s efficiency strategy erodes market share. NVIDIA (NVDA) could see a shift in demand from high-end GPUs for training to more efficient inference chips, though this is speculative. The delay of Gemini 3.5 Pro also risks ceding the frontier model lead to competitors.

These are analytical inferences, not certainties. Actual market reactions will depend on independent validation of Google’s claims and competitive responses.

Alternative Scenarios and Uncertainties

If Gemini 3.5 Pro launches soon with strong performance, the current delay may be seen as a minor setback. Conversely, if Gemini 3.6 Flash fails to deliver meaningful improvements over its predecessor, user adoption could disappoint. Google’s claims about efficiency are self-reported, and independent benchmarks are needed to verify the 65 percent token reduction. The model’s code generation quality, a known pain point, will also be closely scrutinized. Until third-party evaluations emerge, a cautious approach is warranted.

Key Metrics to Watch Next

Investors and analysts should monitor the following indicators: a concrete release date and performance benchmarks for Gemini 3.5 Pro; independent assessments of Gemini 3.6 Flash‘s token cost savings and code quality; and adoption rates of the cybersecurity model by enterprises. Google’s next earnings call may provide insights into cloud AI revenue trends. The next few months will determine whether Google’s efficiency-first strategy is a winning bet or a sign of falling behind in the AI race.

#GoogleAI #Gemini #AIModels #AICosts #Efficiency #Cybersecurity #AIIndustry

Sources

Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.

Found an error? Request a correction →