QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

Qwen3.8-Max-Preview: 2.4 trillion parameters, and open questions

Alibaba teases a giant multimodal model, days after the release of Kimi K3. Behind the announcement, no technical report and no independent benchmark. Here is what we know, and what it changes for local deployment.

Qwen3.8-Max-Preview: 2.4 trillion parameters, and open questions
Short answer. On 19 July 2026, Alibaba presented a preview of Qwen3.8-Max, a multimodal mixture-of-experts model announced at 2.4 trillion parameters, accessible through its API at a reduced rate. The announcement provides no technical report, no model card, no licence and no independent benchmark. This model does not deploy on-premises. For a local platform, the useful range remains models with downloadable weights, such as Qwen 3.6.

What is this about?

On 19 July 2026, Alibaba's Qwen team presented Qwen3.8-Max-Preview, the next iteration of its "Max" range. The model reportedly handles text, images, video and documents. Alibaba claims a size of 2.4 trillion parameters and a position just behind the best models on the market on its own evaluations. The preview is open to developers at ten percent of the standard rate through the Token Plan, and the full version is promised with no timeline.

Every figure calls for caution: this is a commercial announcement, not a scientific publication. No technical report accompanies the preview, and no independent team has been able to measure the model.

A mixture of experts at very large scale

Like the previous "Max" models, the architecture relies on a sparse mixture of experts: the network contains a large number of experts, and a router activates only a small fraction of them for each token. You get the capacity of a giant model at a lower compute cost per request. The open models we deploy, such as DeepSeek V4, Kimi K3 or MiniMax M3, follow the same principle at a documented scale.

The number of active parameters of Qwen3.8-Max has not been communicated. Yet that is the figure that determines the real inference cost, the latency and the memory sizing. Our guide on inference hardware explains why this number drives everything else.

Performance figures to handle with caution

No benchmark table accompanies the preview. The public figures available cover the previous generation, Qwen3-Max, and come from Alibaba's blog. They omit several direct competitors, which rules out an honest comparison.

BenchmarkQwen3-Max score (previous generation)
MMLU-Pro76.0
GPQA71.9
AIME (mathematics)80.0
LiveCodeBench v574.1

Source: Alibaba blog. The SWE-bench figure published at the launch of Qwen3-Max has since been removed from the official blog, which invites reserve on the remaining scores.

An announcement in a tight race

The preview lands days after Moonshot's open-weight launch of Kimi K3. The timing is no accident: Chinese labs multiply frontier-model announcements to occupy the field and attract developers. Teasing a 2.4-trillion-parameter model before any documentation is as much communication as product release.

What we know, what we do not

Known (announced)Unknown (not published)
Name: Qwen3.8-Max-PreviewActive parameters and number of experts
Preview live since 19 July 2026Technical report and model card
Multimodal: text, image, video, documentLicence and availability of the weights
Access at 10% of the rate via the Token PlanIndependent benchmarks and context length

What it changes for local deployment

A 2.4-trillion-parameter model remains out of reach on-premises for nearly every organisation, and the weights are not published anyway. Access goes through Alibaba's API, with the data-location and control questions that raises. Our comparison on-premises or API covers this trade-off.

A local platform is built on models whose specification is published and whose weights can be downloaded: the version is pinned, the hardware is sized to the actual need and the behaviour stays stable over time. Our open-source LLM comparison lists these models, and the Qwen 3.6 page covers the Qwen range that can actually be deployed. A spectacular announcement does not change this criterion: without weights or documentation, a model does not enter an enterprise architecture.

Frequently asked questions

What is Qwen3.8-Max-Preview?

A preview of Alibaba's next flagship model, presented on 19 July 2026. Alibaba announces a multimodal mixture-of-experts model with 2.4 trillion parameters, accessible through its API at a reduced rate. The full version has no release date.

Can Qwen3.8-Max be deployed locally?

No. The weights are not published, no licence has been announced, and a size of 2.4 trillion parameters exceeds the infrastructure of nearly every organisation. Access goes through Alibaba's API, with the data-location questions that raises.

Are the figures announced by Alibaba verified?

No. The preview comes with no technical report, no model card and no independent benchmark. The 2.4-trillion-parameter figure and the positioning against competing models rest solely on Alibaba's own claims.

Which Qwen models can be deployed on-premises?

Qwen models with open weights, such as Qwen 3.6, can be downloaded and run on a single GPU card or a workstation. Their specification is published, which allows precise hardware sizing and a platform that stays stable over time.

Choose a model you can deploy at your site

A call to scope the right open model and the hardware to run it, on your premises.

Book a call

References