Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Key Points
- 3.6 Flash reduces output tokens by ~17%
- 3.5 Flash-Lite: 350 output tokens/s for high throughput
- 3.5 Flash Cyber available via limited CodeMender pilot
Summary
Google releases three new Gemini models tuned for scalable agentic workflows: 3.6 Flash (workhorse with improved efficiency and quality), 3.5 Flash-Lite (high-throughput, low-latency 3.5-class), and 3.5 Flash Cyber (cybersecurity-specialized model integrated with CodeMender for a limited pilot). These aim to reduce token costs, lower latency, and improve reliability across coding, knowledge work, multimodal tasks, and security use cases.
Key Points
-
3.6 Flash
- 17% fewer output tokens vs 3.5 Flash (Artificial Analysis Index) and lower cost per task.
- Pricing: $1.50 / 1M input tokens; $7.50 / 1M output tokens.
- Better coding, knowledge work, and multimodal performance; fewer reasoning steps and tool calls.
- Notable benchmark gains (examples): DeepSWE 49% vs 37%; MLE Bench 63.9% vs 49.7; OSWorld-Verified 83.0% vs 78.4.
- Built-in client-side computer use; Frontier Safety safeguards for CBRN and cyber misuse.
-
3.5 Flash-Lite
- Fastest 3.5-class model at ~350 output tokens/sec (Artificial Analysis).
- Pricing: $0.30 / 1M input tokens; $2.50 / 1M output tokens.
- Optimized for high-throughput, low-latency agentic tasks (search, document processing, pipelines).
- Outperforms prior Flash-Lite generations and in many cases 3.0/3.0x Flash on agentic/coding benchmarks (e.g., Terminal-Bench 54% vs 31%; GDM-MRCR v2 72.2% vs 60.1).
-
3.5 Flash Cyber + CodeMender
- Cybersecurity-focused fine-tune of 3.5 Flash for finding/patching vulnerabilities at scale.
- Integrated into CodeMender multi-agent workflow; competitive on CyberGym.
- Access controlled: limited pilot available only to governments and trusted partners to reduce misuse risk.
-
Availability & next steps
- 3.6 Flash and 3.5 Flash-Lite are available now via the Gemini API (Google AI Studio, Android Studio) and in consumer/enterprise surfaces (Gemini app, Gemini Enterprise); 3.6 also in Antigravity; 3.5 Flash-Lite rolling into Search.
- 3.5 Pro is in partner testing; Gemini 4 pretraining has started.
Practical guidance for engineers
- Choose 3.6 Flash when you need a balance of accuracy and token-cost savings for coding, knowledge, and multimodal agent workflows.
- Use 3.5 Flash-Lite for high-throughput, low-latency pipelines or as a fast subagent in multi-agent systems.
- If you are a verified government or trusted security partner, consider the CodeMender pilot with 3.5 Flash Cyber for vulnerability triage and automated patching tasks.
- Review the model cards and integrate the built-in computer-use tool where available; validate safety and access controls for sensitive or dual-use deployments.