This is incredible — Qwen 3.8 Max is seriously revolutionary.
A new open-source model that beats Claude Fable and Opus 4.8 — despite being several times cheaper. It’s even cheaper than Kimi K3.
Only the second open-source model in history to rank Top 5 in the AI Arena Leaderboard.
Beats Fable comfortably — and look at the massive gap it has ahead of Opus 4.8 — next version 3.9 will very likely beat Opus 5 too:

Open-source has really started to dominate the AI ecosystem. The tide is turning rapidly. Long gone are the days of seeing them as weaker variants of the real deal.
We’ve seen the shocking results from Kimi K3, DeepSeek v4 — and now this.
Don’t be surprised if a few months from now all the top 5 models on any major leaderboard are all open-source.
Qwen 3.8 Max scored 86.6 on Terminal-Bench 2.1, outperforming Claude Fable 5 (84.6) in one of the industry’s toughest evaluations of real-world, agentic terminal coding workflows
And open-source is also becoming contagious — before now the Max model in the Qwen series was never open-source. It was always the lesser models that had their weights made publicly available.
But things have changed completely now — now with Qwen 3.8 even the top-tier Max model is open-source.
AI is getting a lot cheaper and more accessible — both for coding and building AI-powered apps.
10+ days of non-stop coding…
We’ve gotten wild reports of the model being able to sustain 10+ days of autonomous project development — managing hundreds of Git commits while continuously improving software.
Qwen 3.8 Max has been heavily optimized for agentic software development — enabling it to tackle all your most complex engineering tasks from start to finish.
On PaperBench, which measures an AI’s ability to reproduce and understand machine learning research, Qwen 3.8 Max posted an impressive 93.0, comfortably ahead of Claude Opus 4.8 (88.8)
It’s been designed from the ground up for:
- Continuous scientific research
- Multi-hour data science tasks
- Long-horizon planning
- Extended agent workflows
Qwen 3.8 Max achieved 92.6 on GPQA Diamond—a benchmark of graduate-level scientific reasoning—surpassing Claude Opus 4.8 (91.0) while remaining competitive with the strongest frontier models
Perhaps most importantly, Qwen 3.8 Max maintains context remarkably well across lengthy sessions — reducing the looping and repetitive behavior that often appears during large refactoring projects.
Scary for the competition
Major improvements across software engineering and agent benchmarks.

Qwen 3.8 Max outperforms several competing frontier models on Terminal-Bench while remaining highly competitive on software engineering benchmarks like SWE-bench Pro and FrontierSWE.
State-of-the-art self-correcting model
One of the most interesting observations from early testers isn’t reflected in benchmark charts.
Developers report a noticeable improvement in reflective self-correction.
Rather than confidently pushing forward with flawed reasoning, Qwen 3.8 Max often:
- Detects logic errors mid-generation
- Backtracks when it finds faulty assumptions
- Simplifies over-engineered architectures
- Challenges poor software design instead of blindly agreeing
The result is a model that feels more deliberate and less eager to reinforce bad engineering decisions.
Sophisticated architecture for massive parameter size
Behind Qwen 3.8 Max is an enormous 2.4 trillion-parameter Sparse Mixture-of-Experts (MoE) architecture.
That makes it one of the largest publicly disclosed AI models ever built.
Key info:
- 2.4 trillion total parameters
- Sparse Mixture-of-Experts architecture
- Second-largest disclosed AI model, behind Moonshot AI’s 2.8T-parameter Kimi K3
- First Qwen model to exceed one trillion parameters
Unlike dense models, MoE architectures activate only the experts needed for each request, allowing Qwen to deliver frontier-level capability while keeping compute costs significantly lower.
Just as importantly, this is the Qwen team’s first trillion-scale multimodal model.
It can natively understand and reason across:
- Text
- Code
- Images
- Videos
- Complex documents
That makes it suitable for enterprise knowledge management, software engineering, scientific research, and large-scale document analysis.
Qwen 3.8 Max represents one of Alibaba’s most ambitious AI releases yet.
Its biggest differentiators are clear:
- Open weights for the first-ever Qwen Max model
- Frontier-scale 2.4T-parameter architecture
- Strong coding and multi-agent performance
- Long-horizon autonomous reasoning
- Lower deployment costs than many proprietary alternatives
It has massively raised expectations for what developers should expect from an open-weight frontier AI model.