Google Launches Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber as Its Flagship Pro Model Stays Delayed
Editorial Team

Google DeepMind has released three new additions to its Gemini model family on Tuesday, Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite and a specialized cybersecurity model called Gemini 3.5 Flash Cyber, expanding its lineup of fast, low‑cost models even as its more capable Gemini 3.5 Pro remains in limited testing with partners.
The three releases mark Google's second major Gemini update since introducing 3.5 Flash at its I/O conference earlier this year, and reflect a broader industry shift toward cheaper, more efficient models built specifically to power AI agents rather than chase frontier reasoning benchmarks.
Gemini 3.6 Flash Cuts Token Usage While Improving Coding
Gemini 3.6 Flash serves as the centerpiece of the release, positioned as Google's workhorse model for coding, knowledge work and multimodal tasks. According to Google, the model consumes 17 percent fewer output tokens than 3.5 Flash on the Artificial Analysis Index, with token savings reaching as much as 65 percent on certain DeepSWE software engineering benchmarks. The company said the model also requires fewer reasoning steps and tool calls to complete multi‑step workflows, translating directly into lower costs for developers running agentic applications at scale.
On pricing, Gemini 3.6 Flash comes in at 1.50 dollars per million input tokens and 7.50 dollars per million output tokens, notably cheaper than the 9 dollar per million output token rate for 3.5 Flash. Google said the model delivers higher precision with fewer unwanted code edits and reduced execution loops compared to its predecessor, producing more production‑ready code in testing.
Flash‑Lite Becomes the Fastest Model in the Lineup
Gemini 3.5 Flash‑Lite is built for speed and high‑volume use rather than deep reasoning, and Google says it now outpaces even the larger Gemini 3 Flash on agentic software engineering and computer use tasks, despite its smaller size. The model is being positioned as the fastest option across the entire 3.5 Flash family, running at roughly 350 tokens per second, aimed at developers who need rapid responses across large volumes of simple, repetitive tasks.
Both Gemini 3.6 Flash and 3.5 Flash‑Lite are rolling out starting today across the Gemini app, Google AI Studio, Android Studio and the Gemini Enterprise app. Gemini 3.6 Flash is additionally available in Google Antigravity, while 3.5 Flash‑Lite is expanding into Google Search as well.
A Cybersecurity Model Built for Defenders, Not the Public
The most tightly restricted of the three releases is Gemini 3.5 Flash Cyber, a lightweight model built on top of 3.5 Flash and fine‑tuned specifically to find, validate and patch software vulnerabilities. The model operates through CodeMender, Google DeepMind's code security agent, with multiple Flash Cyber subagents working in parallel and combining their findings into a single consolidated vulnerability report.
On the CyberGym benchmark, Flash Cyber scored 83.2 percent, landing within roughly two points of larger competing cybersecurity models while running at a considerably lower cost per token, according to Google. Given the dual‑use nature of vulnerability discovery technology, access to Flash Cyber will initially be limited to governments and trusted partners through a controlled pilot program delivered via CodeMender, with the company saying broader availability will expand gradually over time as it monitors for misuse.
An Intensifying Price War in AI Models
The release lands amid growing pressure across the industry to control ballooning AI compute costs. Google leadership has pointed out that companies are running through their annual token budgets well ahead of schedule, arguing that a mix of smaller, cheaper Flash‑tier models can save large enterprises well over a billion dollars a year in aggregate compute spending, compared with relying primarily on frontier‑scale models for every task.
That efficiency push comes as Google continues to face pressure on the frontier side of its roadmap. Gemini 3.5 Pro, originally teased around Google's I/O conference in May, remains in limited testing with select partners rather than a full public release, leaving competitors room to advance their own top‑tier models in the meantime. Google has said it has already begun what it describes as its most ambitious pretraining run yet for Gemini 4, suggesting the company is looking further ahead even as the 3.5 Pro release timeline remains unclear.
For now, Google's strategy appears focused on winning the efficiency and cost argument at the Flash tier while its frontier model catches up, betting that most real‑world agentic workloads do not require maximum reasoning depth so much as reliable, affordable performance at scale.