📊 Full opportunity report: How Qwen3.8-Max’s AI Numbers Challenge The Competition on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba has officially released Qwen3.8-Max, a 2.4 trillion-parameter AI model with confirmed benchmark results surpassing many competitors. The open weights will be available next week, signaling a significant step in AI model transparency and performance.
Alibaba has officially released detailed benchmark results for Qwen3.8-Max, confirming it as the largest open-weight AI model to date with 2.4 trillion parameters. This development marks a significant milestone in AI performance and transparency, positioning Alibaba’s model as a serious contender against established models like GPT-5.6 and Claude Fable 5, while also offering open access to its weights next week.
Following weeks of speculation and stealth previews, Alibaba revealed the full specifications and benchmark scores for Qwen3.8-Max, a model built on the Qwen3.5 architecture that employs sparse mixture-of-experts techniques. The model features approximately 95 billion active parameters per query within its 2.4 trillion total parameters, and supports multimodal inputs — text, images, and videos — with text output.
In benchmark tests conducted on Alibaba’s own infrastructure, Qwen3.8-Max achieved top scores on several key evaluations. It scored 86.6 on Terminal-Bench 2.1, surpassing Claude Opus 4.8 and Claude Fable 5, and only trailing GPT-5.6 Sol at 88.8. It also led in PaperBench at 93.0, and demonstrated exceptional performance in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. Notably, the model significantly improved its agentic capabilities, with deep software-engineering benchmarks like DeepSWE jumping from 21.6 to 56.6, indicating a major leap in autonomous task execution.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba’s Benchmark Results for AI Leadership
Alibaba’s release and detailed benchmarking of Qwen3.8-Max demonstrate its capability to challenge established AI giants in both size and performance. The model’s high scores on industry-standard benchmarks and its advanced agentic abilities suggest it could influence AI development directions, especially in multimodal and autonomous applications. Furthermore, the open release of the weights next week signifies a shift toward greater transparency and democratization in large-scale AI models, potentially accelerating innovation and deployment across sectors.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Development and Recent Releases
Over the past few weeks, Alibaba's AI efforts have been shrouded in secrecy, with the company initially teasing a model described as 'second only to Fable 5' with 2.4 trillion parameters. The model, later identified as Qwen3.8-Max, was previewed in stealth during the July World AI Conference and appeared on community leaderboards under the alias 'kaleb.' The model's specifications and benchmark results have now been fully disclosed, confirming its status as a major player in the AI landscape.
Earlier, Alibaba's Kimi K3, with 2.8 trillion parameters, briefly caused a stir in the AI market, but Qwen3.8-Max's detailed benchmark results and open weights mark a more substantial and transparent step forward. The company’s strategy of staged announcements and delayed full disclosures reflects a cautious approach to competing in the rapidly evolving AI field.
"The open weights of Qwen3.8-Max will be available next week, marking a milestone in model transparency and accessibility."
— Alibaba spokesperson
high performance GPU for AI development
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Uncertainties Surrounding Model Licensing and Real-World Performance
It remains unclear what the licensing terms for the open weights will be, and whether they will permit broad commercial use. Additionally, while benchmark scores are promising, real-world deployment performance and robustness across diverse tasks are still untested, and the impact of the model’s large size on inference efficiency remains to be seen.

NOVATECH AI Workstation Desktop PC – Intel Core i9-14900K, Liquid Cooling – Machine Learning, Data Science, 3D Rendering, Video Editing, Simulation (RTX 5080 | 64GB RAM | 2TB)
- High-Performance CPU: Intel Core i9-14900K processor
- Powerful GPU: NVIDIA RTX 5080 with 16GB VRAM
- Advanced Cooling System: Liquid cooling for optimal performance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s AI Model Deployment and Community Engagement
Alibaba plans to release the open weights of Qwen3.8-Max next week, enabling researchers and developers to evaluate its capabilities firsthand. The company may also publish further details on licensing and deployment guidelines. Meanwhile, industry observers will monitor how the model performs in practical applications and whether it influences market dynamics or prompts competitors to accelerate their own releases.
As an affiliate, we earn on qualifying purchases.
Key Questions
When will Alibaba release the open weights for Qwen3.8-Max?
The open weights are expected to be available next week, following the benchmark disclosure on August 3, 2023.
How does Qwen3.8-Max compare to other leading models?
In benchmark tests, Qwen3.8-Max outperformed several models like Claude Opus 4.8 and Fable 5 on key evaluations, and was only slightly behind GPT-5.6 in top scores.
What are the main capabilities of Qwen3.8-Max?
The model supports multimodal inputs (text, images, videos), has advanced agentic capabilities, and demonstrates significant improvements in autonomous task execution.
What are the licensing implications for the open weights?
The licensing terms are still unpublished, and it is unclear whether they will allow broad commercial use or impose restrictions.
Will the model’s large size affect its deployment in real-world applications?
Yes, the 2.4 trillion parameters require multi-node data center infrastructure, limiting self-hosting options and raising questions about inference efficiency and accessibility.
Source: ThorstenMeyerAI.com