Skip to main content
Back to News
news/AI Models

Google Launches Gemini 3.6 Flash, Slashes Token Costs by 65%, and Kicks Off Gemini 4 Training

In a late-night triple-drop, Google DeepMind just released three new Gemini models — including one that cuts token consumption by up to 65% on coding tasks — and simultaneously confirmed it has starte

Stefan Trbojevic

Stefan Trbojevic

22 July 20262 min read
LinkedIn
Google Launches Gemini 3.6 Flash, Slashes Token Costs by 65%, and Kicks Off Gemini 4 Training

The takeaway

Google isn't just shipping models — it's reshaping the economics of running AI agents in production, with token costs collapsing and Gemini 4 already in training.

Why it matters for builders

Google isn't just shipping models — it's reshaping the economics of running AI agents in production, with token costs collapsing and Gemini 4 already in training.

Google Launches Gemini 3.6 Flash, Slashes Token Costs by 65%, and Kicks Off Gemini 4 Training

In a late-night triple-drop, Google DeepMind just released three new Gemini models — including one that cuts token consumption by up to 65% on coding tasks — and simultaneously confirmed it has started its "most aggressive pre-training campaign in history" for Gemini 4. The message is unmistakable: Google is going all-in on making AI agents cheaper, faster, and smarter.

Three Models, One Night

On July 22, 2026, Google unveiled:

  • Gemini 3.6 Flash — a next-gen workhorse that dramatically reduces cost and token usage
  • Gemini 3.5 Flash-Lite — an ultra-fast, ultra-cheap variant hitting 350 tokens per second
  • Gemini 3.5 Flash Cyber — a cybersecurity specialist built exclusively for vulnerability detection and remediation

Gemini 3.6 Flash: The Cost-Killer

The headline feature of 3.6 Flash is brute-force efficiency. According to the Artificial Analysis Index, it uses 17% fewer output tokens than its predecessor 3.5 Flash. On the DeepSWE coding benchmark, savings hit 65% — fewer detours, fewer redundant reasoning steps, fewer unnecessary tool calls per task.

Pricing reflects the efficiency gains: $1.50 per million input tokens and $7.50 per million output tokens — cheaper than 3.5 Flash despite stronger benchmark scores. Key performance jumps include:

Benchmark 3.5 Flash 3.6 Flash
DeepSWE (coding) 37% 49%
MLE Bench (ML research) 49.7% 63.9%
OSWorld-Verified (computer use) 78.4% 83%
GDPVal-AA v2 (knowledge tasks) +70 pts

The model also shines at multi-agent orchestration, demonstrated through complex code migration and real-time 3D workflow generation via Gemini Canvas.

Flash-Lite: Small Model, Big Punch

Gemini 3.5 Flash-Lite hits 350 tokens per second at a staggering $0.30/M input, $2.50/M output — purpose-built for high-volume scenarios like bulk document processing and agentic search.

Most impressively, this lightweight model outperforms its larger sibling Gemini 3 Flash on both SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). The intended architecture: 3.6 Flash as the "master brain" handling planning, with a fleet of Flash-Lite workers executing at scale.

Flash Cyber: Security-First AI

Gemini 3.5 Flash Cyber tackles a painful reality: AI finds vulnerabilities faster than existing systems can patch them. Integrated into Google's CodeMender platform, it has already set a new state-of-the-art on the CyberGym cybersecurity benchmark — using a lightweight model where competitors need heavyweight ones. It is not publicly available for now.

Gemini 4: The Real Bombshell

Beyond the three launches, Google confirmed it has internally kicked off its most aggressive pre-training run in history, targeting Gemini 4. Industry observers note that, following Google's typical ~6-month training cycle, Gemini 4 could arrive as early as late 2026.

Key takeaway: Google isn't just shipping models — it's reshaping the economics of running AI agents in production. With token costs collapsing, a lightweight model beating its larger predecessor, and Gemini 4 already in the oven, the cost of building on Google's AI stack is dropping faster than almost anyone predicted.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

22 July 2026

Updated

22 July 2026

Sources

Source links pending editorial review.

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.