Briefly
- Google launched Gemini 3.6 Flash and three.5 Flash-Lite in the present day, with higher effectivity and decrease prices than 3.5 Flash—however Gemini 3.5 Professional, promised at I/O 2026 for June supply, stays in testing after falling quick on coding internally.
- 3.6 Flash makes use of 17% fewer output tokens than 3.5 Flash whereas dropping the output worth from $9 to $7.50 per million tokens, making it cheaper to run AI brokers at scale.
- Google confirmed it has begun pre-training for Gemini 4, which it calls “our most bold pre-training run but.”
Google launched three new AI fashions in the present day: Gemini 3.6 Flash, 3.5 Flash-Lite, and three.5 Flash Cyber. That wasn’t what most individuals anticipated.
After unveiling Gemini 3.5 Flash at Google I/O 2026 in Could and promising a Professional model inside a month, Google quietly missed its personal deadline. Gemini 3.5 Professional was held again as a result of it fell wanting inner targets, per Bloomberg, significantly on coding duties. A late-June try to repair it by updating the coaching knowledge—the huge datasets a mannequin learns from—produced disappointing outcomes. Alphabet inventory fell roughly 4.4% on the report, erasing an estimated $200 billion in market cap in a single session.
The final Professional-tier mannequin Google shipped was Gemini 3’s successor, Gemini 3.1 Professional, again in February.
The Flash collection is Google’s line of speed-optimized fashions—quick, cost-effective, and constructed for AI brokers, that are packages that function semi-autonomously to deal with duties like managing paperwork, processing knowledge pipelines, or searching the net and not using a human clicking by every step. Professional fashions are the heavy lifters: slower, pricier, and constructed for complicated reasoning the place uncooked energy issues greater than pace.
What every AI mannequin does—and who it is for
Gemini 3.6 Flash is the primary launch. It makes use of 17% fewer output tokens—tokens being the fundamental unit AI processes, roughly three-quarters of a phrase—than 3.5 Flash, per the Synthetic Evaluation Index. It is also cheaper: $1.50 per million enter tokens and $7.50 per million output tokens, down from $9 on the output facet for 3.5 Flash. For companies working brokers at scale, that distinction compounds quick.
On benchmarks—standardized checks that rating AI by share of duties accomplished accurately—3.6 Flash hit 49% on DeepSWE v1.1, which checks long-horizon software program engineering like constructing and debugging full codebases, versus 37% for 3.5 Flash. On MLE-Bench, a machine studying engineering check, it scored 63.9% versus 49.7%. It topped the desk on OSWorld-Verified—a check the place the AI takes management of a pc display to finish actual duties—at 83.0%, forward of Claude Sonnet 5 (81.2%) and GPT-5.6 Luna (72.6%).

Rivals in the identical class nonetheless lead elsewhere: GPT-5.6 Luna scores 67% on DeepSWE and 84.7% on Terminal-Bench 2.1, which checks agentic terminal coding. Claude Sonnet 5 tops information work on GDPval-AA v2—a benchmark scored on an Elo score scale like chess, the place greater numbers imply higher real-world activity efficiency—at 1607 versus 3.6 Flash’s 1421.
We tried the mannequin for coding and the outcomes had been… underwhelming to say the least.
Our easy coding check ended up with an unusable file. The HTML was not correctly formatted, and components weren’t rendered accurately. Subsequent makes an attempt to vibe code a solution to clear up the problems weren’t profitable.

We requested Deepseek to show the primary mannequin into one thing playable by merely fixing the bugs. It recognized 11 bugs and carried out 8 key fixes, which resulted in an honest sport.

Deepseek’s small tweaks mounted the sport, which implies, Gemini’s core pondering was right, however the particulars and inaccuracies made the end result unuseful. Put together for lengthy vibe coding classes with an affordable but poor performing mannequin in case you faux to make use of the mannequin for that.

The second mannequin launched by Google, Gemini 3.5 Flash-Lite, is constructed purely for quantity: 350 output tokens per second at $0.30/million enter and $2.50/million output. It is geared toward high-throughput pipelines—assume doc processing at huge scale or agentic search programs—and outperforms the older 3 Flash on key coding duties, together with Terminal-Bench 2.1 (54% vs. 31%), regardless of being considerably cheaper.
It may be an incredible session compactor (analyzing lengthy classes and extracting the important thing components so your agent doesn’t collapse with noise) for these counting on Hermes and Openclaw.
The third mannequin, Gemini 3.5 Flash Cyber, will not be publicly out there. Google is proscribing it to governments and vetted companions who want to seek out and repair software program vulnerabilities—a dual-use functionality the corporate shouldn’t be comfy releasing broadly.
In the meantime, Google’s DeepMind crew is already transferring on. Google confirmed within the official announcement that it has began “our most bold pre-training run but, for Gemini 4,” and the crew is already hyping it up.
Pre-training is the foundational part the place a mannequin learns from huge datasets earlier than task-specific fine-tuning begins—which means Gemini 4 is being constructed, not deliberate.
We now have began our most bold pre-training run but, for Gemini 4, and are excited by the progress : )
— Logan Kilpatrick (@OfficialLoganK) July 21, 2026
Each 3.6 Flash and three.5 Flash-Lite are stay in the present day within the Gemini app, Google AI Studio, and by way of the API. Gemini 3.5 Professional will ship, per Google, “as quickly because it’s prepared,” every time that’s.
Every day Debrief E-newsletter
Begin each day with the highest information tales proper now, plus unique options, a podcast, movies and extra.
