Overview
Qwen3.8-Max is Alibaba Cloud's flagship model, served from Alibaba Cloud Model Studio as qwen3.8-max with a dated snapshot qwen3.8-max-0902. It reached general availability on August 3, 2026, having been previewed on July 19, 2026 at the World AI Conference in Shanghai as qwen3.8-max-preview โ a preview that shipped a headline parameter count and nothing else: no model card, no active-parameter figure, no scores.
Two things about it are worth more than any benchmark row.
The first is scale and sparsity. Qwen3.8-Max is a 2.4 trillion parameter Mixture-of-Experts model that activates 95 billion parameters per token. Alibaba describes it as built on the Qwen3.5 architectural foundation and scaled up. Where Qwen3.7-Max published no architecture at all, this generation publishes the whole config, because you can download it.
The second is the licence, and it inverts what this page said three months ago. The Qwen Max line is no longer closed. The base checkpoint Qwen/Qwen3.8-2.4T-A95B went up on huggingface.co/Qwen on August 12, 2026, with an FP8 build alongside it, under a custom Qwen3.8-Max License. Qwen3.7-Max had no weights anywhere and Alibaba's flagship had looked like a permanent API product; that is no longer the description.
Read the licence before you plan around it, though. It is not Apache 2.0 โ see Limitations โ and the downloadable checkpoint is not the model the API serves. Alibaba's own model card is explicit: "Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support." The open weights are text-only and thinking-required, with a 262,144-token native window. The hosted model adds vision and video input, a non-thinking path, and the full 1M window.
The third change from Qwen3.7-Max is modality. The previous flagship's documentation answered "Can I upload images or documents as input?" with "These models accept text only," and routed visual work to Qwen3-VL. Model Studio now lists qwen3.8-max under image and video understanding as well as text generation.
Capabilities
- Native vision and video input: The hosted model takes text, image and video. Alibaba's launch material frames the 1M window in terms of what it holds โ "hundred-page documents, full television series, or 100-hour livestreams" โ rather than in tokens. This is new to the Max line as of 3.8.
- Sparse MoE at 2.4T: 95B active parameters per token out of 2.4T total. Serving cost tracks the active count, which is why a 2.4T model prices below the previous 3.7 generation rather than above it.
- Selectable reasoning depth:
xhigh(the default) for tasks demanding thorough analysis,mediumto balance accuracy against speed,lowfor cost- and latency-sensitive work.preserve_thinkingis enabled by default across workloads. - 1M token context: The hosted API serves the full million; output is capped at 131,072 tokens.
- Context caching: Cached input reads at 10% of the standard input rate on both the Singapore and Beijing cards.
- Long-horizon autonomy: Alibaba's launch post reports a single unattended run of roughly 125 hours in which the model wrote about 7,600 lines of code, took over 1,100 actions and ran 33 rounds of GPU training. It is a vendor anecdote, not a benchmark, but it is the capability the model is being sold on.
- Harness portability: Alibaba claims comparable results across QwenWork, Claude Code, Codex, OpenClaw and Hermes โ a claim about scaffold-independence that is straightforward to test on your own harness.
- Downloadable base weights:
Qwen3.8-2.4T-A95Band an FP8 build, for teams that need to self-host and can meet the licence terms and the hardware bill.
Technical Specifications
- Hosted model IDs:
qwen3.8-max(rolling alias),qwen3.8-max-0902(snapshot, September 2, 2026),qwen3.8-max-preview(July 19, 2026) - Total parameters: 2.4 trillion
- Active parameters: 95 billion per token
- Architecture: sparse MoE; the open checkpoint's config declares 512 experts with 11 activated per layer (10 routed + 1 shared), in a hybrid stack alternating Gated DeltaNet blocks with gated attention
- Context window: 1,000,000 tokens hosted; 262,144 native on the open checkpoint, extensible to 1,010,000
- Max output: 131,072 tokens
- Modalities: text, image and video in, text out (hosted); text only on the open weights
- Reasoning:
xhigh/medium/low;preserve_thinkingon by default - Open weights:
Qwen/Qwen3.8-2.4T-A95BandQwen/Qwen3.8-2.4T-A95B-FP8, published August 12, 2026 - Licence: Qwen3.8-Max License โ free for most commercial use, with revenue- and MAU-gated conditions
- Platforms: Alibaba Cloud Model Studio and DashScope; QwenWork; Hugging Face and ModelScope for the weights
Alibaba has still not published a training-token count or a knowledge cutoff for Qwen3.8-Max. Those fields are absent from this page rather than estimated.
The open-weight line in this generation
| Repository | Published | Type | Context | Licence |
|---|---|---|---|---|
| Qwen3.8-2.4T-A95B | 2026-08-12 | MoE, 2.4T total / 95B active, text-only | 262,144 (โ1,010,000) | Qwen3.8-Max License |
| Qwen3.8-27B | 2026-08-14 | Dense 27B, vision-language | 262,144 (โ1M via YaRN) | Apache 2.0 |
The smaller model kept Apache 2.0; the flagship did not. If permissive licensing is the hard requirement rather than open weights as such, the 27B is the model to build on. Qwen3.8-Flash-Next, a 125B MoE with 6B active, followed on August 26 under qwen-community-1.0 โ we covered its local-inference story in Qwen3.8-Flash-Next Runs Locally in 75 GB of RAM.
The prior release: Qwen3.7-Max
Qwen3.7-Max, launched May 20, 2026 at the Alibaba Cloud Summit in Hangzhou, remains available on Model Studio as qwen3.7-max with the qwen3.7-max-2026-05-20 and qwen3.7-max-2026-06-08 snapshots. It was closed-weight and API-only, text-only, with hybrid thinking on by default and a 1M token context window. Its Singapore rate card is $2.50 input and $7.50 output per 1M tokens โ above the 3.8 flagship on both sides. The lower-cost qwen3.7-plus sibling is still the tiered option: $0.40 / $1.60 per 1M up to 256K input tokens, rising to $1.20 / $4.80 beyond it. There is no qwen3.8-plus; the cheap slot in this generation is qwen3.8-flash.
Use Cases
- Long-horizon coding agents: Multi-day unattended runs with tool use and self-correction are the workload Alibaba built and benchmarked for. The 1M window plus
xhighreasoning is the intended configuration. - Video and document understanding at length: Native video input against a million-token window suits transcript-plus-frames analysis of long recordings โ meetings, broadcasts, livestream archives โ without a separate vision model in the pipeline.
- Whole-repository reasoning: A 1M window holds a substantial codebase in a single request, removing much of the retrieval scaffolding.
- Sovereign or air-gapped deployment: Newly possible on this line.
Qwen3.8-2.4T-A95Bcan be self-hosted under vLLM, SGLang or comparable runtimes โ subject to the licence gates and to finding the cluster. - Cost-sensitive frontier work: At $2.00 / $6.00 it undercuts the Western frontier tier substantially while claiming comparable scores. Whether the scores hold on your tasks is the whole question.
- Multilingual enterprise workloads: The Qwen series' long-standing strength in Chinese and English, now with a self-hosted fallback.
Performance / Benchmarks
Alibaba published a first-party evaluation table with the launch post, "Qwen3.8-Max: A New Bar for Coding and Cowork" (Alibaba Cloud Community, August 3, 2026). All figures below are Alibaba-reported, measured by Alibaba, at launch.
A methodological note before the numbers: Alibaba renders its benchmark tables as images, and the qwen.ai blog serves no static text at all. The scores below are transcribed from the launch material and cross-checked against independent write-ups of it; treat the second decimal with more suspicion than you would a machine-readable table. Comparator scores are reproduced as Alibaba stated them, under the names Alibaba used.
| Benchmark | Qwen3.8-Max | Comparators cited by Alibaba |
|---|---|---|
| Terminal Bench 2.1 | 86.6 | GPT-5.6 Sol 88.8; Claude Opus 4.8 84.6; Fable 5 84.6 |
| SWE-bench Pro | 67.7 | Fable 5 80.0; Opus 4.8 69.2; GPT-5.6 Sol 64.6 |
| PaperBench | 93.0 | GPT-5.6 Sol 90.5; Fable 5 88.8; Opus 4.8 80.3 |
| GPQA Diamond | 92.6 | GPT-5.6 Sol 94.1; Fable 5 92.6; Opus 4.8 92.0 |
| IFBench | 82.8 | GPT-5.6 Sol 72.7; Fable 5 63.5; Opus 4.8 62.2 |
| HLE | 43.6 | Fable 5 53.3; GPT-5.6 Sol 47.2; Opus 4.8 45.7 |
| OSWorld-Verified | 86.1 | Fable 5 85.0; GPT-5.6 Sol Max 83.2; Gemini 3.1 Pro 76.2 |
| MathVision | 95.2 | โ |
| LogicVista | 91.9 | โ |
The shape of that table is more informative than any single row. Qwen3.8-Max leads on instruction-following (IFBench, by twenty points over the Claude entries), on research reproduction (PaperBench) and on computer use (OSWorld-Verified), and trails on the two hardest reasoning and software-engineering rows โ HLE and SWE-bench Pro โ by wide margins. It is not a uniform frontier claim and Alibaba does not present it as one.
Four caveats travel with these numbers:
- They are vendor-run and unreplicated. No independent evaluation of Qwen3.8-Max had been published at the time of writing.
- Most coding rows used the Claude Code harness, not a neutral scaffold. Harness choice moves agentic coding scores materially.
- Several benchmarks are Alibaba's own constructions and cannot be audited by a third party โ including the E-Commerce Bench simulation Alibaba highlights, where it reports a ยฅ416,252 closing balance, a 4.16x return, and a 38% margin over second-place GLM-5.2.
- The comparator names are Alibaba's. Where Alibaba cites "Claude Opus 4.8," "Fable 5" or "GPT-5.6 Sol," the configuration behind that label is not independently documented in the post.
For the September snapshot, Alibaba reports qwen3.8-max-0902 at 1,691 on Code Arena WebDev, and large jumps on two coding evaluations โ TerminalBench 3.0 from 11.3 to 29.0 and ProgramBench from 10.5 to 28.0. Those are first-party figures on a post-training refresh, and the low absolute values are the interesting part: both benchmarks are hard enough that the frontier is still under 30.
Limitations
- Open weight is not Apache 2.0. The Qwen3.8-Max License is broadly permissive but carries two gates: display the model name in your UI above 100M MAU or US$20M monthly revenue, and negotiate a separate licence if you run a Model-as-a-Service or "AI Work Assistant" business above US$50M revenue in any twelve months. Legal review is not optional for anyone near those thresholds.
- The open checkpoint is a different model from the API. No vision, no video, no non-thinking mode, and 262,144 native context rather than 1M. Benchmarks measured on the hosted model do not transfer to a self-hosted deployment.
- Self-hosting 2.4T is a datacenter problem. Sparse activation cuts compute per token, not storage. The full-precision weights are a multi-node deployment; FP8 helps but does not make this a single-server model.
- No published knowledge cutoff and no training-token count. Ground time-sensitive queries with retrieval or tool use.
- Benchmarks are vendor-run, harness-favourable and partly internal. See the four caveats above.
- Thinking is on by default at the deepest setting.
xhighis the default reasoning effort, so latency and output-token spend both run high unless you explicitly step down tomediumorlow. - Regional price divergence, and two currencies. Singapore is quoted natively in USD, Beijing natively in CNY. Cost models must account for which region serves the request and for the exchange rate Alibaba applies when restating one card in the other's currency.
- Preview-era documentation was thin, and some of it persists. The July preview shipped without a model card; third-party summaries written in that window still circulate with no active-parameter figure and no scores.
Pricing & Access
Alibaba quotes each region in its own currency, and the two rate cards are not conversions of one another. The International (Singapore) card is natively in US dollars. The Chinese mainland (North China 2, Beijing) card is natively in Chinese yuan. Each pricing page restates the other region's figures in its own currency, which is where the long decimals come from. Prices are per 1M tokens.
International (Singapore) โ quoted in USD
Source: Alibaba Cloud Model Studio pricing.
| Model | Input | Cached input | Output |
|---|---|---|---|
qwen3.8-max | $2.00 | $0.20 | $6.00 |
qwen3.8-flash | $0.15 | โ | $0.47 |
qwen3.7-max (prior flagship) | $2.50 | $0.25 | $7.50 |
qwen3.8-max is flat across its entire 1M window โ no tiering. Cached input reads at 10% of the standard rate.
Chinese mainland โ North China 2 (Beijing) โ quoted in CNY
Source: ้ฟ้ไบ็พ็ผๆจกๅ่ฎก่ดน ("Alibaba Cloud Model Studio model pricing"), section ๅๅ2๏ผๅไบฌ๏ผ.
| Model | Input | Output |
|---|---|---|
qwen3.8-max | ยฅ12 | ยฅ36 |
qwen3.8-flash | ยฅ0.8 | ยฅ2.7 |
Alibaba's English page restates the mainland qwen3.8-max row as $1.65 / $4.951. Those are converted figures, not the vendor's native quote โ the round yuan numbers are.
Access options:
- Alibaba Cloud Model Studio โ the primary managed API, and where the snapshots are pinned
- DashScope โ Alibaba's model-serving platform
- QwenWork โ Alibaba's workplace agent platform, in public beta since the August 3 launch
- Hugging Face and ModelScope โ
Qwen3.8-2.4T-A95Band its FP8 build, under the Qwen3.8-Max License - Third-party aggregators โ resell the hosted API
Ecosystem & Tools
- Qwen3.8-Max: A New Bar for Coding and Cowork โ Alibaba's launch post and the primary source for the benchmark table
- Alibaba unveils Qwen3.8-Max โ the press-room announcement, August 3, 2026
- Alibaba Cloud Model Studio โ supported models, capabilities and snapshot IDs
- Model Studio pricing (English) โ per-region rate cards
- ็พ็ผๆจกๅ่ฎก่ดน (Chinese) โ the mainland rate card as Alibaba natively quotes it, in CNY
- Qwen/Qwen3.8-2.4T-A95B โ the open base checkpoint, its config, and the licence text
- Qwen on Hugging Face โ the full open-weight catalogue, including Qwen3.8-27B under Apache 2.0
- QwenLM on GitHub โ inference recipes and model repositories
- ModelScope โ Alibaba's model community
Community & Resources
- Supported models and capabilities overview
- Text generation models on Model Studio
- Qwen
- Our write-up of the smaller sibling: Qwen3.8-Flash-Next Runs Locally in 75 GB of RAM
- Compare with DeepSeek V4, Kimi K3, GLM-5.3, and Tencent Hy3