Google launched Gemini 3.7 Flash on Thursday, a low-cost coding-and-agents mannequin now typically out there.
OpenAI previewed GPT-5.6 Sol Ultrafast, a Cerebras-powered tier working its most succesful mannequin as much as 14× sooner.
The velocity race has shifted from uncooked intelligence to real-time brokers, however solely Google’s mannequin reaches each developer at this time.
Google and OpenAI each pushed the identical message at this time: AI is now quick sufficient to really feel spectacular for many who use AI brokers, every firm asserting extremely quick fashions.
The 2 launches are constructed in another way, and just one is definitely in your fingers at this time.
Myriad: When will OpenAI launch GPT-6? Click on to make your prediction.
Google shipped Gemini 3.7 Flash, its newest mannequin tuned for software program coding and autonomous enterprise workflows. OpenAI opened a restricted preview of GPT-5.6 Sol Ultrafast, a brand new service tier that runs its most succesful mannequin at as much as 750 output tokens per second.
At the moment we’re introducing Gemini 3.7 Flash, our most clever workhorse mannequin but for coding and brokers.
This mannequin brings substantial beneficial properties throughout software program engineering, net growth, and sophisticated data work.
Gemini 3.7 Flash is a general-availability mannequin. It takes as much as one million enter tokens (roughly 750,000 phrases) and returns 64,000, dealing with textual content, photos, video, audio and PDFs, and it may name instruments and management a pc. Google is pitching it as a budget mind for autonomous methods that plan duties and end multi-step jobs with much less human assist.
The mannequin isn’t sacrificing high quality for velocity. It’s each extra succesful and extra environment friendly, having the ability to full our take a look at coding activity in 2 minutes and 13 seconds whereas the newest Flash mannequin took greater than 5 minutes. The standard hole between the 2 can also be noticeable.
OpenAI’s Ultrafast is not a brand new mannequin. It is GPT-5.6 Sol—the identical mannequin OpenAI used an AI purple workforce to harden in opposition to prompt-injection assaults earlier than launch—on a sooner observe, powered by chipmaker Cerebras. It’s round 14 instances sooner than GPT-5.6 Sol’s personal customary velocity.
Previewing Ultrafast mode: GPT-5.6 Sol at as much as 14x the velocity.
Launching first within the OpenAI API to a choose group of consumers with expanded entry to extra companies as capability grows. pic.twitter.com/a5dleofiDJ
— OpenAI (@OpenAI) August 13, 2026
Cerebras’ wafer-scale chips generate as much as 750 tokens a second, about 560 phrases, quick sufficient {that a} voice agent can suppose mid-call.
The numbers that matter
Google’s personal benchmark sheet places Gemini 3.7 Flash forward of Claude Sonnet 5, GPT-5.6 Terra, and others on 11 of 18 examined classes, together with a prime Code Enviornment web-dev rating of 1,588 Elo and 30.4% on AutomationBench for enterprise workflows. That’s all primarily based on Google’s methodology, so deal with the lead as the corporate’s declare.
As any Gemini Flash, the mannequin can also be low cost. At 75 cents per million enter tokens and $3.75 per million output tokens by means of year-end, it is half of Gemini 3.6 Flash’s unique fee. That intro worth expires December 31, then doubles to $1.50 and $7.50, which continues to be low cost for a Google mannequin.
OpenAI hasn’t revealed head-to-head scores for Ultrafast past buyer quotes. Jane Avenue AI engineer John Crepezzi stated in OpenAI’s announcement that Cerebras’ velocity “allows alternative ways of utilizing the fashions.” Podium product lead Courtland Lykins known as it “invaluable in our voice stack,” saying the velocity “utterly modifications the decision expertise.” The tier is invite-only for now.
The velocity push lands because the labs pivot from “who’s smartest” to “who’s quick sufficient for brokers.” Google’s timing is pointed. Its flagship Gemini 3.5 Professional continues to be lacking, with no launch date given, three weeks after 3.6 Flash and days after a DeepMind management reshuffle that moved Demis Hassabis apart for deputy Koray Kavukcuoglu. OpenAI, in the meantime, is renting Cerebras’ velocity fairly than ready by itself stack.
Gemini 3.7 Flash is stay now in additional than 160 nations; GPT-5.6 Sol Ultrafast continues to be invite-only.
Every day Debrief Publication
Begin each day with the highest information tales proper now, plus unique options, a podcast, movies and extra.