Google adds cheaper Gemini models as AI cost race intensifies
Alphabet is rolling out three Gemini models focused on lower costs, faster workloads and cybersecurity ahead of its earnings report.
By Theo Nakamura · Staff Writer
· 3 min read
Alphabet is expanding its Gemini AI lineup with three new models, giving investors a fresh read on how Google plans to compete on price, speed and specialized use cases. The launch matters because AI models are expensive to run, and lower usage costs can shape how quickly companies adopt them.
Google said Tuesday it is releasing Gemini 3.5 Flash Cyber, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The update arrives a day before Alphabet reports earnings, with the company trying to show progress in artificial intelligence as rivals in the U.S. and China push ahead.
Gemini 3.5 Flash Cyber is built to identify and repair software security flaws, according to Google. The company said access will start through a limited pilot for governments and trusted partners, rather than a broad public release.
The cybersecurity model is also priced below Google’s larger models on a per-token basis, the company said. A token is a small unit of data that an AI model processes, often a piece of a word or other input. Lower token costs can reduce the bill for developers and companies that run many AI requests.
Google pushes the cost angle
Gemini 3.6 Flash is aimed at coding, multimodal work and knowledge tasks, according to Google. Multimodal means the model can handle more than one type of input, such as text and images.
Google said Gemini 3.6 Flash uses up to 17% fewer tokens than the prior version and costs less per token. For high-volume customers, that kind of efficiency can matter because AI expenses rise with repeated model calls, longer prompts and larger workloads.
Gemini 3.5 Flash-Lite is designed as the fastest and lowest-cost model in Google’s 3.5 family, according to the company. Google said it is meant for smaller tasks and high-volume work, including jobs inside larger AI-agent systems. An AI agent is software that uses a model to complete tasks across several steps.
Artificial Analysis data cited by CNBC shows Gemini Flash already costs less than comparable models from Anthropic, OpenAI and Chinese competitors. According to the firm, Gemini 3.6 Flash is cheaper per task than GPT-5.6 Terra Max, Kimi K3 and Qwen 3.7 Max. Artificial Analysis said Gemini 3.5 Flash-Lite costs far less than those models.
Competition is moving fast
The cybersecurity model gives Google a more direct answer to Anthropic, which CNBC reported has built an early edge in automated code defense. Anthropic’s lead has made security-focused AI a more visible battleground for major model companies.
Chinese AI developers are also applying pressure. Moonshot AI’s Kimi K3 saw enough demand that the company restricted new subscriptions and API access because of capacity limits, according to a post from Kimi. API access lets outside software connect to a model.
Alibaba has teased Qwen 3.8 Max, saying it ranks behind only Anthropic’s Fable 5 on overall performance, according to the company’s post cited by CNBC.
That demand points to a practical limit in the AI race: companies need enough computing power to serve models after they build them. Google may have an advantage from its custom chips, cloud systems and ability to design hardware and software together, although CNBC reported the company has dealt with capacity constraints too.
CNBC also reported that Google is developing a specialized chip intended to run Gemini up to 10 times more efficiently. A Google Cloud spokesperson told CNBC that the company is “constantly researching and experimenting with new innovations” and said not every project reaches production.
The spokesperson added that Google’s hardware and software are co-designed so its systems are optimized for real-world workloads, according to CNBC.
Google is also giving more detail on future releases. The company said Gemini 3.5 Pro is being tested with partners before wider availability, and it has started its largest pre-training run yet for Gemini 4. Pre-training is the early phase where a model learns patterns from large datasets before it is adapted for specific uses.
This story draws on original reporting from CNBC.