Gemini 3.8 Flash is absolutely insane

Google’s new Gemini 3.8 Flash is absolutely insane.

One of the fastest models in the world right now — and gives you a high amount of intelligence for incredibly low cost.

3.8 Flash is only second behind Gemini 3.5 Flash Lite in speed in the Artificial Analysis leaderboard:

It even beat much bigger and costlier models like GPT-5.6 Sol in multiple coding benchmarks.

Massive upgrades in agentic and long-horizon coding

Gemini 3.8 Flash is much better at long-horizon coding — working of big coding tasks that need many many steps — or hours long debugging.

On DeepSWE v1.1 Gemini 3.8 Flash scored 73.7%. Gemini 3.7 Flash scored 65.3%.

That’s a huge 8.4 percentage point jump.

And it matched to several bigger models — Claude Opus 5 scored 74%. GPT-5.6 Sol scored 72.7% — Gemini 3.8 Flash came within 0.3 points of Opus 5 — and beat GPT-5.6 Sol.

Terminal-Bench 2.1 is even more impressive.

Gemini 3.8 Flash scored 89.4% compared with 85.8% for 3.7 Flash.

It also beat Claude Opus 5 at 89.1% and GPT-5.6 Sol at 88.8%.

You’ll realize just how much of a big deal this is when you see the massive difference in cost.

Unbelievably high intelligence per cost

Gemini 3.8 Flash is launching at a incredible price of $0.75 per million input tokens and $3.75 per million output tokens.

That makes its coding results even more impressive.

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. GPT-5.6 Sol costs $4 per million input tokens and $20 per million output tokens.

So Gemini is more than 80% cheaper for both input and output tokens

And this difference will really show up as you spend hours coding with agents.

An agent can read dozens of files, call tools, run code, think about the result and try again…

All of that burns tokens.

Gemini 3.8 Flash gives developers a lot of intelligence for very little money — exactly what you want when agents start working for hours instead of answering one prompt.

Unbelievably fast

Usually the smartest models take a long time to think and come up with results.

But with Gemini 3.8 Flash you’re getting very high intelligence at blazing fast speeds.

It kept the speed of 3.7 Flash while making a big jump in coding and reasoning.

That gives it an insane intelligence-to-speed ratio.

And speed matters even more for agents.

If an agent needs 30 tool calls to finish something then every delay adds up.

A model that is both smart and fast can completely change how these products feel.

Scarily good at hacking

Google also made Gemini 3.8 Flash Cyber.

This version focuses on defensive cybersecurity — and its results are wild.

On CyberGym it scored 86.2% for autonomous vulnerability discovery. GPT-5.5-Cyber scored 85.6% while GPT-5.6 Sol scored 83.6%.

Google also tested the model across 20 programming languages.

It successfully found vulnerabilities 71% of the time — up from 58.9% for Gemini 3.7 Flash.

Google says its Chrome Security team also got 2.6 times more correct vulnerability patches from 3.8 Flash Cyber than from the best much-larger commercial models it tested.

That power comes with obvious risks.

So Google isn’t giving everyone unrestricted use. The Cyber model sits behind its Fairwind Program for trusted defenders.

No more Minimal thinking

One smaller change says a lot about this model.

Google killed minimal thinking.

Gemini 3.8 Flash only gives developers three thinking levels:

Low. Medium. High.

Even on low — it thinks.

3.8 Flash can spend more tokens than its predecessor on hard problems. It takes more reasoning steps and uses tools more often.

That could well explain its long-horizon coding gains.

The AI race isn’t just about who has the smartest model anymore.

It’s becoming a race to deliver the most intelligence for the least money and time.

Gemini 3.8 Flash makes a very strong case for that future.



Leave a Comment

Your email address will not be published. Required fields are marked *