AI

Google Launches Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber

Google has introduced three new Gemini Flash models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, delivering improved token efficiency, record speed, and specialized cybersecurity capabilities for building production AI agents.

Google released three Gemini Flash models on July 22, 2026: Gemini 3.6 Flash improves coding and token efficiency at a lower price; Gemini 3.5 Flash-Lite delivers 350 tokens per second for high-throughput agentic tasks; and Gemini 3.5 Flash Cyber is a specialized cybersecurity model deployed via a limited-access pilot. The updates focus on reducing token waste and latency while raising benchmark scores across coding, computer use, and knowledge work.

Google dropped a trio of new Gemini Flash models on July 22, 2026, each tuned for a different slice of the agent-building stack. The flagship Gemini 3.6 Flash delivers higher coding and reasoning accuracy with measurably lower token consumption. Gemini 3.5 Flash-Lite becomes the fastest model in the 3.5 family, hitting 350 output tokens per second. And Gemini 3.5 Flash Cyber brings a specialized, cost-efficient capability for finding and fixing code vulnerabilities at scale.

The common thread across all three is a sharper focus on token efficiency and latency. Google's developer feedback from the earlier 3.5 Flash generation made one thing clear: when you are running multi-step agentic workflows, every redundant output token and every unnecessary tool call burns money. The new releases attack that problem from different angles.

Gemini 3.6 Flash shows a 17 percent reduction in output tokens compared to 3.5 Flash on the Artificial Analysis Index, and in benchmarks like DeepSWE, output token usage drops by as much as 65 percent. That efficiency comes with performance gains, not sacrifices. On DeepSWE itself, 3.6 Flash scores 49 percent versus 37 percent for 3.5 Flash. On MLE Bench, it reaches 63.9 percent, up from 49.7 percent. Computer-use capabilities also improve, reaching 83.0 percent on OSWorld-Verified versus 78.4 percent. Pricing sits at 1.50permillioninputtokensand1.50 per million input tokens and 7.50 per million output tokens, lower than 3.5 Flash while delivering better results.

Gemini 3.5 Flash-Lite targets high-throughput, latency-sensitive tasks. With 350 output tokens per second, it is the fastest model in the 3.5 series. Priced at 0.30permillioninputtokensand0.30 per million input tokens and 2.50 per million output tokens, it dramatically undercuts previous Flash-Lite generations while posting significant quality improvements. On Terminal-Bench 2.1, it hits 54 percent versus 31 percent for 3.1 Flash-Lite. On long-context benchmarks like GDM-MRCR v2, it reaches 72.2 percent. Even against the older 3 Flash, 3.5 Flash-Lite wins on several agentic and coding evals, including SWE-Bench Pro (54.2 percent vs. 49.6 percent).

Gemini 3.5 Flash Cyber is built on 3.5 Flash and fine-tuned specifically for cybersecurity workflows. Paired with Google's CodeMender agent, it can find, validate, and patch vulnerabilities at a fraction of the cost of larger models. Google is deploying it through a limited-access pilot program for governments and trusted partners, deliberately restricting broader availability due to the dual-use nature of the technology.

All three models are available immediately in the Gemini API and Google AI Studio. Gemini 3.6 Flash is also live in Vertex AI. Meanwhile, Google confirmed that Gemini 3.5 Pro is in partner testing and that the team has already kicked off pre-training for Gemini 4, described as the company's most ambitious training run to date.