AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Qwen3.8-Max’s Latest AI Stats: A New Contender In The Race on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of its Qwen3.8-Max AI model, revealing benchmark scores and confirming open weights will ship next week. The model features 2.4 trillion parameters and shows competitive performance in several benchmarks, signaling a significant development in large language models.

Alibaba has confirmed the broad availability of Qwen3.8-Max, a large language model with 2.4 trillion parameters. The company published detailed benchmark scores and announced that open weights will be shipped next week, marking a significant milestone in its AI development.

On August 3, Alibaba unveiled the full benchmark table for Qwen3.8-Max, which had been withheld since its preview in July. The model is built on the Qwen3.5 architecture, uses sparse mixture-of-experts, and supports multimodal input—text, image, and video—with text output. The active-parameter count is approximately 95 billion per query, with the total parameters at 2.4 trillion.

The benchmark results show Qwen3.8-Max outperforming several competitors in key tests: it scored 86.6 on Terminal-Bench 2.1, surpassing Claude Fable 5 and only behind GPT-5.6 Sol. It achieved top scores on PaperBench (93.0) and demonstrated strong multimodal and agentic capabilities, notably improving in long-horizon tasks and agentic execution relative to its predecessor. Alibaba also confirmed that the model can reproduce research results and outperform previous models on specific tasks like AIME24, with notable gains in agentic benchmarks.

The company clarified that the 2.4 trillion parameters are a theoretical total, with about 95 billion active parameters per query, and that the open weights, due next week, will be a checkpoint suitable for deployment on high-memory hardware. The open weights will be available under unspecified licensing, with the model’s deployment primarily via API, compatible with OpenAI and DashScope interfaces.

At a glance
updateWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba officially released detailed benchmark results and confirmed the upcoming release of open weights for its Qwen3.8-Max model, marking a notable advancement in AI model capabilities.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's AI Benchmark Performance

The release of detailed benchmark scores and upcoming open weights signifies a major step for Alibaba in the large language model landscape. The model's strong performance in key benchmarks like Terminal-Bench and PaperBench demonstrates its competitiveness against established players like OpenAI and Anthropic. Additionally, the open weights will enable broader access for developers and researchers, potentially accelerating innovation and deployment in multimodal and agentic AI applications.

However, the model's performance varies across benchmarks, with notable gaps in deep software engineering tasks, such as SWE-bench Pro, where it trails behind Fable 5. The distinction between the 2.4 trillion total parameters and the approximately 95 billion active parameters per query highlights the model's sparse architecture, which influences its practical deployment and scalability. Overall, this development signals a significant push by Alibaba into the large-scale AI arena, with implications for competition, open-access AI development, and future capabilities of multimodal models.

/Modern GPU Programming with Rust and CUDA 13: Mastering Parallel Computing, GPU Acceleration, Memory Optimization, AI Systems, and High-Performance Application Development (Learning Express Series)

/Modern GPU Programming with Rust and CUDA 13: Mastering Parallel Computing, GPU Acceleration, Memory Optimization, AI Systems, and High-Performance Application Development (Learning Express Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Recent Developments in Alibaba's AI Strategy

Alibaba's AI journey has been marked by strategic teasers and stealthy previews, culminating in the July preview of Qwen3.8-Max, which was initially introduced via an anonymous model called 'kaleb' at the Code Arena leaderboard. The company confirmed its identity during the World AI Conference in Shanghai on July 19, after a series of teasers and limited previews. Prior to this, Alibaba's AI models had been less publicly detailed, with the company focusing on building a competitive large language model ecosystem.

The model's announcement follows a period of intense activity in the AI industry, characterized by high-profile launches such as Moonshot's Kimi K3 and the emergence of other large models. Alibaba's approach involved a staged reveal, starting with a stealth preview, then limited access via a paid endpoint, and now full benchmark disclosure and open weights, reflecting a strategic effort to build credibility and showcase capabilities.

Historically, Alibaba's open models have shipped under Apache 2.0 licenses, but the upcoming release of the 2.4 trillion parameter checkpoint remains uncertain in licensing terms. The company emphasizes that the open weights will be a multi-node datacenter artifact, not suitable for individual hosting, but the 27B checkpoint will be designed for local deployment.

"Qwen3.8-Max demonstrates the company's commitment to advancing multimodal and agentic AI capabilities, with benchmark results confirming its competitive edge."

— Alibaba spokesperson

Unresolved Questions About Open Weights and Licensing

While Alibaba confirmed that open weights will be shipped next week, details about the licensing terms remain unpublished. It is unclear whether the open weights will be under a permissive license like Apache 2.0 or a more restrictive one, which could impact adoption. Additionally, the performance of the 27B checkpoint—derived from the 2.4T model—on various benchmarks and real-world tasks is still to be evaluated once released.

It is also uncertain how the model's agentic capabilities will hold up in practical deployments, especially given the gaps in deep software engineering benchmarks. The long-term impact of the model's architecture and sparse mixture-of-experts design on scalability and efficiency remains to be seen.

Upcoming Release of Open Weights and Benchmark Evaluations

Next week, Alibaba plans to release the open weights of Qwen3.8-Max, enabling broader testing and deployment. Researchers and developers will likely scrutinize the model's performance across benchmarks and real-world applications, especially in multimodal and agentic tasks. Meanwhile, the company may publish further details about licensing and usage terms.

Industry watchers will be monitoring how the open weights influence the competitive landscape, especially against models like GPT-5.6 and Claude Fable 5. The model's long-term impact on AI capabilities and open-access development will become clearer as more tests and deployments occur.

Key Questions

What are the key specifications of Alibaba's Qwen3.8-Max?

The model has 2.4 trillion total parameters, with approximately 95 billion active parameters per query. It supports multimodal input—text, image, and video—and is built on a sparse mixture-of-experts architecture based on Qwen3.5.

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be shipped next week, with Alibaba confirming their availability but not yet publishing the licensing details.

How does Qwen3.8-Max compare to other large models?

Benchmark scores show Qwen3.8-Max outperforming several models like Claude Fable 5 in some tests, with scores of 86.6 on Terminal-Bench 2.1, but trailing behind GPT-5.6 Sol at 88.8. Its multimodal and agentic capabilities are notably strong, especially in long-horizon tasks.

What are the potential implications of this release?

The release could accelerate AI development and deployment, especially in multimodal and agentic applications, but licensing and practical deployment details will influence its adoption and impact.

Source: ThorstenMeyerAI.com

You May Also Like

The Compute Reckoning: Anthropic Finally Admits What Customers Suspected for Ten Months

Anthropic confirms that its recent customer experience issues were due to insufficient compute capacity, now addressed by a major SpaceX deal and other investments.

The Ultimate Guide To Security Layers For AI Agent Infrastructure

A new security proxy for MCP servers introduces per-tool allowlists, identity verification, and audit logs to enhance AI agent infrastructure security.

“This is going to be a niche device” – Analysts react to the $1,000+ Steam Machine price reveal

Experts say the new Steam Machine’s high price limits its appeal, positioning it as a niche device within the gaming market.

The Labor Displacement Data: What Q1-Q2 2026 Actually Shows

New data from early 2026 shows AI-driven layoffs are focused on specific cohorts, with overall labor market stability despite significant structural shifts.