Google releases Gemini Flash upgrades while Pro model remains delayed
The faster, cheaper Gemini models target AI agents, but Bloomberg reported that Gemini 3.5 Pro is still in testing after missing internal goals.
By Sofia Marchetti · Columnist
· 3 min read
Google rolled out three Gemini AI models, giving developers cheaper and faster tools while leaving its more powerful Pro model on the sidelines. For Alphabet investors, the missing Pro release matters because the market is watching whether Google can keep pace in high-end AI, where product delays can hit sentiment fast.
Google announced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, according to a company blog post. The release follows Google’s May unveiling of Gemini 3.5 Flash at I/O 2026, where the company said a Pro version would arrive within a month.
That Pro model has not shipped. Bloomberg reported that Gemini 3.5 Pro remained in testing after falling short of Google’s internal goals, especially on coding tasks. Bloomberg also reported that a late-June effort to improve the model by refreshing its training data did not deliver the results Google wanted.
Alphabet shares fell about 4.4% after the Bloomberg report, wiping out an estimated $200 billion in market value in one session. The last Pro-tier Gemini release was Gemini 3.1 Pro in February.
Flash is about speed and cost
Google’s Flash models are built for speed and lower operating costs. That makes them useful for AI agents, which are software systems that can carry out multi-step tasks with less human input, such as moving through websites, processing files or handling data workflows.
Pro models sit at the other end of the lineup. They are meant for harder reasoning work, where higher cost and slower response times may be acceptable if the model performs better on complex tasks.
Gemini 3.6 Flash is the main launch. According to the Artificial Analysis Index, it uses 17% fewer output tokens than Gemini 3.5 Flash. Tokens are the small chunks of text that AI systems read and generate, roughly comparable to pieces of words. Fewer output tokens can lower costs when a company runs a model many times across many users or agents.
Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. The prior 3.5 Flash model cost $9 per million output tokens, according to the figures cited.
Where the new model scored well
On DeepSWE v1.1, a benchmark that tests long software-engineering tasks such as building and fixing codebases, Gemini 3.6 Flash scored 49%, according to Artificial Analysis. Gemini 3.5 Flash scored 37% on the same test.
On MLE-Bench, which measures machine-learning engineering performance, Gemini 3.6 Flash scored 63.9%, compared with 49.7% for Gemini 3.5 Flash. On OSWorld-Verified, a benchmark where an AI model controls a computer screen to finish real tasks, Gemini 3.6 Flash scored 83.0%, ahead of Claude Sonnet 5 at 81.2% and GPT-5.6 Luna at 72.6%, according to the cited benchmark data.
Rivals still led in other areas. GPT-5.6 Luna scored 67% on DeepSWE and 84.7% on Terminal-Bench 2.1, which tests terminal-based coding agents. Claude Sonnet 5 led on GDPval-AA v2, a knowledge-work benchmark scored using an Elo-style rating system, with 1607 versus 1421 for Gemini 3.6 Flash.
Google also said it has begun pre-training Gemini 4, calling it “our most ambitious pre-training run yet.” Pre-training is the phase where an AI model learns patterns from large datasets before later tuning and testing.
This story draws on original reporting from Decrypt.