How Qwen3.8-Max’s AI Numbers Challenge The Competition

📊 Full opportunity report: How Qwen3.8-Max’s AI Numbers Challenge The Competition on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has officially released Qwen3.8-Max, a 2.4 trillion-parameter AI model with confirmed benchmark results surpassing many competitors. The open weights will be available next week, signaling a significant step in AI model transparency and performance.

Alibaba has officially released detailed benchmark results for Qwen3.8-Max, confirming it as the largest open-weight AI model to date with 2.4 trillion parameters. This development marks a significant milestone in AI performance and transparency, positioning Alibaba’s model as a serious contender against established models like GPT-5.6 and Claude Fable 5, while also offering open access to its weights next week.

Following weeks of speculation and stealth previews, Alibaba revealed the full specifications and benchmark scores for Qwen3.8-Max, a model built on the Qwen3.5 architecture that employs sparse mixture-of-experts techniques. The model features approximately 95 billion active parameters per query within its 2.4 trillion total parameters, and supports multimodal inputs — text, images, and videos — with text output.

In benchmark tests conducted on Alibaba’s own infrastructure, Qwen3.8-Max achieved top scores on several key evaluations. It scored 86.6 on Terminal-Bench 2.1, surpassing Claude Opus 4.8 and Claude Fable 5, and only trailing GPT-5.6 Sol at 88.8. It also led in PaperBench at 93.0, and demonstrated exceptional performance in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. Notably, the model significantly improved its agentic capabilities, with deep software-engineering benchmarks like DeepSWE jumping from 21.6 to 56.6, indicating a major leap in autonomous task execution.

At a glance
updateWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba announced the broad availability of Qwen3.8-Max, revealing detailed benchmark results that challenge existing AI models in both size and performance.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Benchmark Results for AI Leadership

Alibaba’s release and detailed benchmarking of Qwen3.8-Max demonstrate its capability to challenge established AI giants in both size and performance. The model’s high scores on industry-standard benchmarks and its advanced agentic abilities suggest it could influence AI development directions, especially in multimodal and autonomous applications. Furthermore, the open release of the weights next week signifies a shift toward greater transparency and democratization in large-scale AI models, potentially accelerating innovation and deployment across sectors.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Development and Recent Releases

Over the past few weeks, Alibaba's AI efforts have been shrouded in secrecy, with the company initially teasing a model described as 'second only to Fable 5' with 2.4 trillion parameters. The model, later identified as Qwen3.8-Max, was previewed in stealth during the July World AI Conference and appeared on community leaderboards under the alias 'kaleb.' The model's specifications and benchmark results have now been fully disclosed, confirming its status as a major player in the AI landscape.

Earlier, Alibaba's Kimi K3, with 2.8 trillion parameters, briefly caused a stir in the AI market, but Qwen3.8-Max's detailed benchmark results and open weights mark a more substantial and transparent step forward. The company’s strategy of staged announcements and delayed full disclosures reflects a cautious approach to competing in the rapidly evolving AI field.

"The open weights of Qwen3.8-Max will be available next week, marking a milestone in model transparency and accessibility."

— Alibaba spokesperson

Amazon

high performance GPU for AI development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties Surrounding Model Licensing and Real-World Performance

It remains unclear what the licensing terms for the open weights will be, and whether they will permit broad commercial use. Additionally, while benchmark scores are promising, real-world deployment performance and robustness across diverse tasks are still untested, and the impact of the model’s large size on inference efficiency remains to be seen.

NOVATECH AI Workstation Desktop PC – Intel Core i9-14900K, Liquid Cooling – Machine Learning, Data Science, 3D Rendering, Video Editing, Simulation (RTX 5080 | 64GB RAM | 2TB)

NOVATECH AI Workstation Desktop PC – Intel Core i9-14900K, Liquid Cooling – Machine Learning, Data Science, 3D Rendering, Video Editing, Simulation (RTX 5080 | 64GB RAM | 2TB)

  • High-Performance CPU: Intel Core i9-14900K processor
  • Powerful GPU: NVIDIA RTX 5080 with 16GB VRAM
  • Advanced Cooling System: Liquid cooling for optimal performance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s AI Model Deployment and Community Engagement

Alibaba plans to release the open weights of Qwen3.8-Max next week, enabling researchers and developers to evaluate its capabilities firsthand. The company may also publish further details on licensing and deployment guidelines. Meanwhile, industry observers will monitor how the model performs in practical applications and whether it influences market dynamics or prompts competitors to accelerate their own releases.

Amazon

multimodal AI input devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are expected to be available next week, following the benchmark disclosure on August 3, 2023.

How does Qwen3.8-Max compare to other leading models?

In benchmark tests, Qwen3.8-Max outperformed several models like Claude Opus 4.8 and Fable 5 on key evaluations, and was only slightly behind GPT-5.6 in top scores.

What are the main capabilities of Qwen3.8-Max?

The model supports multimodal inputs (text, images, videos), has advanced agentic capabilities, and demonstrates significant improvements in autonomous task execution.

What are the licensing implications for the open weights?

The licensing terms are still unpublished, and it is unclear whether they will allow broad commercial use or impose restrictions.

Will the model’s large size affect its deployment in real-world applications?

Yes, the 2.4 trillion parameters require multi-node data center infrastructure, limiting self-hosting options and raising questions about inference efficiency and accessibility.

Source: ThorstenMeyerAI.com

You May Also Like

The Atlas. What the framework is.

The Post-Labor Transition Atlas offers an empirical, structural framework analyzing AI’s impact on labor markets, highlighting heterogeneity and policy implications.

The Channel Move: Anthropic, Wall Street, and the Acquisition of the Real Economy

Anthropic partners with Blackstone, H&F, Goldman Sachs, and General Atlantic in a $1.5B joint venture to embed AI into thousands of portfolio companies, transforming enterprise AI deployment.

A Frontier AI Model Just Went Dark for 18 Days. The Kill-Switch Is Real Now.

An advanced AI model was globally disabled for 18 days following US government orders, marking a shift in AI regulation and security protocols.

Data: The One Thing You Can’t Rent

As AI models face data scarcity, industry shifts toward fenced, verified, and proprietary data sources, making data the new chokepoint and competitive edge.