Alibaba on Monday published benchmark scores it had withheld on July 19 for Qwen3.8-Max, its 2.4 trillion-parameter model. The launch post says Qwen3.8-Max has 95 billion active parameters. Many outlets have described it as a sparse mixture-of-experts design, which routes each query to a fraction of its parameters so serving it costs far less than its total size implies. Alibaba also set next week as the release window for the model weights.
What Changed
- Alibaba published a full benchmark table for Qwen3.8-Max on Monday, two weeks after claiming on July 19 that the model was "second only to Fable 5" without releasing any scores.
- The model carries 2.4 trillion total parameters with 95 billion active, about seven times the parameter count of Qwen3.5, which Alibaba released in February.
- Open weights are due next week on Hugging Face and ModelScope, alongside a smaller Qwen3.8-27B. Alibaba has still not stated which license will govern them.
- The scores come from Alibaba's own internal runs. The Decoder reported that independent verification is pending, and the only third-party placements available are crowdsourced Arena.AI rankings.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
What Alibaba disclosed
On July 19, Alibaba called Qwen3.8-Max “second only to Fable 5,” according to an Implicator report based on South China Morning Post’s preview coverage. The report identified three missing disclosures: the activated-parameter count, the license and the weights date. Monday’s release supplied the count and date. Alibaba has not stated which license will govern the weights it plans to publish on Hugging Face and ModelScope.
The model has 2.4 trillion total parameters, about seven times as many as Qwen3.5, which Alibaba released in February. Its context window holds one million tokens, or roughly 750,000 words per query, according to the South China Morning Post. Alibaba said it will also release a smaller Qwen3.8-27B next week. Qwen3.8-Max will be the first Max-class Qwen model with downloadable weights.
The company benchmark table
Alibaba’s table put Qwen3.8-Max at 93.0 on PaperBench, the highest result among the models it listed. On Terminal Bench 2.1, Qwen scored 86.6, behind GPT-5.6 Sol at 88.8. The Decoder noted that these figures came from Alibaba’s internal runs and that independent verification is pending.
Crowdsourced Arena.AI results provide the only third-party placements available at publication. Qwen3.8-Max ranked fifth in Text Arena, second in Vision Arena and fourth in Frontend Code Arena. Alibaba selected those placements for its launch materials.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Price and investor response
Alibaba priced Qwen3.8-Max at $2 per million input tokens and $6 per million output tokens. Hong Kong-listed shares closed Monday up 7%, their largest gain in nearly a month, Bloomberg reported.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Vey-Sern Ling, managing director at Union Bancaire Privée, told Bloomberg, “Many investors continue to underestimate Chinese AI models because of US chip restrictions or general skepticism. In reality, the gap is probably much closer, and narrowing fast. Alibaba’s Qwen 3.8 is another proof point, following Kimi K3.”
Independent testing
The Kimi K3 model Ling cited was released by Moonshot with open weights on July 27. The model has 2.8 trillion parameters, but subsequent independent testing found that it fell well short of top Western models in cyber capabilities and complex math, according to The Decoder.
Andrew Yoon, a technical staff member at CivAI, told The Deep View, “Claims that recent Chinese models match or beat the US frontier are overstating cherry-picked benchmark results. This isn’t to say that the models are weak. Models like Qwen 3.8 Max, Kimi K3, and GLM 5.2 are all very capable, but they continue to lag the US frontier by a significant margin.”
Frequently Asked Questions
What did Alibaba actually disclose on Monday that it had not disclosed before?
On July 19 Alibaba called Qwen3.8-Max "second only to Fable 5" and published no benchmark scores. An Implicator report that day identified three missing disclosures: the activated-parameter count, the license, and the weights date. Monday's release supplied the activated-parameter count, 95 billion, and set next week as the weights window. The license remains unstated.
How large is Qwen3.8-Max, and how much of it runs at once?
The model has 2.4 trillion total parameters, about seven times as many as Qwen3.5, which Alibaba released in February. Forbes Middle East and TNGlobal described it as a sparse mixture-of-experts design, which routes each query to a fraction of those parameters. Alibaba's launch post puts the active count at 95 billion. Its context window holds one million tokens.
How did Qwen3.8-Max score on the benchmarks Alibaba published?
Alibaba's table put the model at 93.0 on PaperBench, the highest result among the models it listed, and 86.6 on Terminal Bench 2.1, behind GPT-5.6 Sol at 88.8. On crowdsourced Arena.AI leaderboards it ranked fifth in Text Arena, second in Vision Arena and fourth in Frontend Code Arena.
Can the benchmark numbers be trusted?
Not independently, not yet. The Decoder reported that the figures came from Alibaba's internal runs and that independent verification is pending. The Arena.AI placements are third-party, but Alibaba selected which ones to publicize. Andrew Yoon of CivAI told The Deep View that claims Chinese models match the US frontier are "overstating cherry-picked benchmark results."
What does Qwen3.8-Max cost, and how did the market react?
Alibaba priced the model at $2 per million input tokens and $6 per million output tokens. Hong Kong-listed shares closed Monday up 7%, their largest gain in nearly a month, Bloomberg reported. Jefferies analysts wrote that Alibaba's full stack of capabilities stands out.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR