Google DeepMind Launches Gemini 3.6 Flash and 3.5 Flash-Lite to Power Faster, Cheaper AI Agents

Google DeepMind Launches Gemini 3.6 Flash and 3.5 Flash-Lite to Power Faster, Cheaper AI Agents

New Gemini models promise sharper coding, multimodal performance and lower costs for developers building large-scale AI agent systems.

Google DeepMind has released two new AI models — Gemini 3.6 Flash and Gemini 3.5 Flash-Lite — designed to make AI agents faster and cheaper to run at scale. Both models became generally available on 21 July 2026.

The company has also announced a third model, Gemini 3.5 Flash Cyber, a security-focused variant aimed at finding and fixing software vulnerabilities. Access to that one is restricted to selected trusted partners for now.

Gemini 3.6 Flash: The New Workhorse

Gemini 3.6 Flash replaces Gemini 3.5 Flash as Google’s main mid-tier model. Google describes it as the new “workhorse” for coding, knowledge work, multimodal tasks and what the company calls agentic workflows — systems where an AI plans and carries out a series of steps using tools and external services, rather than just answering a single question.

The headline efficiency claim is striking. According to Google’s Artificial Analysis Index and supporting technical data, Gemini 3.6 Flash uses about 17% fewer output tokens than its predecessor for comparable tasks. On specific coding benchmarks — including Datacurve’s DeepSWE test — output token reductions of up to 65% have been reported. Fewer tokens means less processing time and lower bills.

Meanwhile, the model supports an input context window of around one million tokens and can generate up to 64,000 output tokens in a single interaction. That makes it capable of working through long documents, complex multi-step reasoning chains and large codebases in one go.

Pricing has also shifted. Google is billing Gemini 3.6 Flash at around $1.50 per million input tokens — that’s roughly £1.15 at current exchange rates — and around $7.50 per million output tokens (around £5.75). The output price is down from the $9 per million charged for Gemini 3.5 Flash. Worth noting: Google’s pricing includes “thinking” tokens — the internal reasoning steps the model works through — even when those steps aren’t shown in the final answer.

Google’s own tweet described Gemini 3.6 Flash as delivering higher-quality work at “the exact same cost” as 3.5 Flash. But external pricing data tells a slightly different story: the official per-token output price has actually fallen. The “same cost” framing appears to refer to overall workload cost for equivalent tasks — a marketing characterisation rather than a like-for-like tariff comparison.

Flash-Lite: Built for Speed and Volume

Gemini 3.5 Flash-Lite is aimed at a different problem: handling enormous volumes of requests quickly and cheaply. Google says it’s suited to agentic search, document processing, content classification and everyday tasks that need to run at scale — think millions of queries per day.

Independent technical reporting based on Google’s briefings puts its throughput at around 350 tokens per second, making it one of Google’s most efficient models for real-time applications. Benchmark performance is described as approaching that of frontier models from roughly a year ago, at a fraction of the price.

Indicative pricing from developer-focused coverage places Gemini 3.5 Flash-Lite at around $0.30 per million input tokens (about £0.23) and $2.50 per million output tokens (about £1.92). Those figures come from public reporting rather than official UK price lists and should be treated as subject to change.

Google has already deployed Gemini 3.5 Flash-Lite inside Google Search itself.

Gemini 3.5 Flash Cyber: A Security-Focused Variant

Gemini 3.5 Flash Cyber is a specialised model built around cybersecurity applications. Unlike its sibling releases, it is not publicly available: Google has restricted access to a group of selected trusted partners, which the company has not fully disclosed. The model is designed to assist with identifying and remediating software vulnerabilities — tasks that sit at the more sensitive end of AI capability, given the potential for misuse if such a tool were made widely accessible. Google has not published benchmark data or pricing for Gemini 3.5 Flash Cyber, and no general availability date has been announced. Its restricted rollout reflects a broader industry caution around releasing security-oriented AI tools into the open market before adequate safeguards are in place.

What the Models Can Do

Both Gemini 3.6 Flash and 3.5 Flash-Lite are multimodal. They can process text, images, audio, video and PDFs within a single API call. Both support function calling, code execution and multi-step agent flows — the building blocks for AI systems that can browse the web, write and run code, or process documents without human intervention at each step.

Developers can access Gemini 3.6 Flash through Google AI Studio, the Gemini API, Android Studio and Google Antigravity — Google’s environment for building agentic systems. A limited free tier is available through Google AI Studio for prototyping, with metered pricing for production use.

Industry Reaction: Progress, But Questions Remain

Technology analysts broadly welcome the efficiency gains but describe them as incremental rather than transformational. Google’s rapid deprecation of Gemini 3.5 Flash — only recently highlighted at the company’s I/O developer conference — underlines just how fast the AI model landscape is moving.

Some observers caution that lower per-token prices don’t automatically mean lower overall bills. If usage volumes rise — as they tend to when costs fall — or if sophisticated agentic workflows consume more “thinking tokens,” the savings can erode quickly.

Jeffery Tang, a developer advocate at Google DeepMind, said: “Gemini 3.6 Flash is designed to handle the most demanding agentic workloads while keeping costs predictable for teams building at scale.”

Data-protection and digital-rights advocates have raised broader concerns. Cheaper, higher-throughput models could encourage more extensive processing of user-generated content, they argue, without equivalent improvements in transparency or safeguards.

Workers in coding, content and back-office roles face a mixed picture. Lower-cost AI tools create new opportunities — new products, new services, new workflows — but also put pressure on tasks that can now be automated more cheaply than before.

What This Means for Kent Residents

For Kent businesses and developers already building on Google Cloud or the Gemini API, Gemini 3.6 Flash and 3.5 Flash-Lite are available now and offer lower costs per task than the models they replace — useful for anyone running document processing, coding assistance or customer-facing chatbots. If Gemini 3.5 Flash-Lite is now powering Google Search, as Google states, Kent residents using Google’s search products may notice subtly sharper, faster answers to everyday queries, though any improvement will be global rather than local. Public bodies such as Kent County Council or NHS Kent and Medway ICB considering AI tools built on these models would need to ensure any use of personal data complies with UK GDPR and the Data Protection Act 2018 before deployment.

Source: @GoogleDeepMind

Google DeepMind Launches Gemini 3.6 Flash and 3.5 Flash-Lite to Power Faster, Cheaper AI Agents Quiz

5 questions