News
Google Ships Three New Gemini Models, Leaves Flagship Pro Version Waiting
- By John K. Waters
- 07/21/2026
Google on Tuesday released three new additions to its Gemini model lineup, a token-efficient workhorse model, a low-cost speed-focused variant, and a specialized cybersecurity model, while confirming that its long-teased flagship Gemini 3.5 Pro remains in partner testing with no firm release date.
The releases, announced in a post on the company's official blog written by Tulsee Doshi, senior director of product management for Gemini, land one day ahead of parent company Alphabet Inc.'s quarterly earnings report and roughly two months after Google introduced Gemini 3.5 Flash at its I/O developer conference in May.
Three Models, One Tier
All three new models sit in Google's Flash tier, which the company positions for speed, cost efficiency and high-volume agentic workloads rather than maximum reasoning depth.
Gemini 3.6 Flash is described by Google as its new workhorse model, built to improve coding, knowledge work and multimodal performance. According to Google's own figures, cited from the Artificial Analysis Index, the model uses 17 percent fewer output tokens than Gemini 3.5 Flash, with the company reporting reductions of up to 65 percent on the DeepSWE coding benchmark from Datacurve. Google also says the model takes fewer reasoning steps and tool calls to complete multi-step workflows. Pricing for 3.6 Flash is set at $1.50 per million input tokens and $7.50 per million output tokens, down from $9 per million output tokens for its predecessor.
Gemini 3.5 Flash-Lite is pitched as the fastest and most cost-effective model in the 3.5 generation, running at 350 output tokens per second by Google's measurement and priced at $0.30 per million input tokens and $2.50 per million output tokens. Google says the model outperforms the company's own Gemini 3 Flash on some agentic and coding evaluations, including the SWE-Bench Pro and OSWorld-Verified benchmarks.
Both 3.6 Flash and 3.5 Flash-Lite are available immediately through the Gemini app, the Gemini API via Google AI Studio and Android Studio, and Google's Gemini Enterprise platform. Flash-Lite is also rolling out inside Google Search.
A Cybersecurity Model With Guardrails on Access
The third release, Gemini 3.5 Flash Cyber, is built on the 3.5 Flash architecture and fine-tuned specifically to detect, validate and patch software vulnerabilities. It powers
CodeMender, Google's automated code security agent, which runs multiple Flash Cyber instances together to produce a single vulnerability report.
In a companion post on the Google DeepMind blog, the company said the model outperformed its own mainline 3.5 Flash and 3.6 Flash models on an internal evaluation built by its Big Sleep security research team, which targets vulnerabilities in complex codebases such as Chrome and Safari. Google also reported that in production testing against the V8 JavaScript engine, Flash Cyber found more confirmed unique issues than either its own mainline Flash model or a competing model from Anthropic. Those figures come from Google's own published evaluations and have not been independently verified.
Because vulnerability-finding tools carry inherent dual-use risk, Google is not releasing Flash Cyber broadly. Access will be limited to governments and what the company calls trusted partners through CodeMender, in what it describes as a limited-access pilot program.
The Model Google Didn't Ship
The releases notably did not include Gemini 3.5 Pro, the higher-capability model Google teased alongside the May launch of 3.5 Flash, when the company said it was already in internal use and would roll out "next month."
Bloomberg reported last week that Google has faced internal delays getting 3.5 Pro to meet its performance targets, according to TechCrunch's account of that reporting. In Tuesday's blog post, Google said only that 3.5 Pro "is currently testing with partners and we plan to make it broadly available as soon as it's ready."
Google DeepMind product lead Logan Kilpatrick, speaking Tuesday, said the company hopes the model will "land soon," and confirmed the team has begun what Google's blog post separately called its "most ambitious pre-training run yet" for Gemini 4.
Timing and Competitive Backdrop
The Tuesday release lands a day before Alphabet reports quarterly earnings, and follows a stretch in which Google has faced pressure to show consistent progress on its AI roadmap amid competition from OpenAI, Anthropic and a fast-improving field of Chinese model developers. CNBC reported that Alphabet chief executive Sundar Pichai said earlier this year that using a mix of Flash-tier models could save enterprise customers more than $1 billion annually, a figure the network attributed to Pichai rather than Tuesday's release materials.
What it Means for Pure AI Readers
For developers and IT buyers evaluating Gemini for production use, the substance of Tuesday's announcement is less about raw capability and more about unit economics.
Google's pitch for 3.6 Flash and 3.5 Flash-Lite centers on token efficiency and lower per-task cost for agentic workloads, where a model may make dozens of tool calls or reasoning steps to complete a single job. If Google's own efficiency figures hold up under independent testing, teams running high-volume agent pipelines, document processing or search-heavy workloads stand to see meaningful cost reduction without switching providers or re-architecting prompts.
The gated rollout of Flash Cyber is worth watching closely. Google is explicitly framing the model as dual-use technology and restricting it to government and vetted partner access through CodeMender rather than the open API. That is a more conservative distribution approach than Google has taken with most Gemini releases, and it signals the company sees real risk in putting a fast, cheap vulnerability-finding model into general circulation. Security teams at regulated organizations should not expect self-service access to this model anytime soon, and should instead watch for expanded pilot eligibility criteria from Google.
The absence of 3.5 Pro is the more consequential story for enterprise buyers weighing Gemini against Anthropic's and OpenAI's top-tier models for complex reasoning tasks. A second consecutive delay, following the missed "next month" timeline from May, suggests Google is prioritizing shipping cost-efficient Flash-tier improvements while its flagship reasoning model remains unfinished. Organizations with active or planned migrations to Gemini for high-complexity workloads should factor that uncertainty into procurement timelines rather than assume Pro-tier parity with competitors is imminent. The Gemini 4 pretraining disclosure suggests Google's longer-term roadmap is intact, but it offers no near-term relief for buyers who need frontier-class reasoning today.
About the Author
John K. Waters is the editor in chief of a number of Converge360.com sites, with a focus on high-end development, AI and future tech. He's been writing about cutting-edge technologies and culture of Silicon Valley for more than two decades, and he's written more than a dozen books. He also co-scripted the documentary film Silicon Valley: A 100 Year Renaissance, which aired on PBS. He can be reached at [email protected].