---
source: 'https://howaiworks.ai/models/qwen'
section: models
title: Qwen3.8-Max
description: >-
  Alibaba's flagship Qwen model, released August 3, 2026. A 2.4T-parameter MoE
  with 95B active, 1M token context, text/image/video input — and open weights.
tags:
  - Qwen
  - Alibaba
  - Language Model
  - Large Language Model
  - MoE
  - Open Weights
  - Multimodal
  - Agentic AI
  - Long Context
  - Latest
category: Language Models
developer: Alibaba Cloud
developerWebsite: 'https://www.alibabacloud.com/'
modelType: Multimodal Language Model
releaseDate: '2026-08-03'
lastUpdated: '2026-09-06'
license: 'Qwen3.8-Max License (open weights, revenue-gated terms)'
contextWindow: 1M tokens
inputPrice: $2.00
outputPrice: $6.00
officialWebsite: 'https://qwen.ai/'
docsPage: 'https://www.alibabacloud.com/help/en/model-studio/models'
pricingPage: 'https://www.alibabacloud.com/help/en/model-studio/model-pricing'
---

# Qwen3.8-Max

> Alibaba's flagship Qwen model, released August 3, 2026. A 2.4T-parameter MoE with 95B active, 1M token context, text/image/video input — and open weights.

## Overview

Qwen3.8-Max is Alibaba Cloud's flagship model, served from Alibaba Cloud Model Studio as `qwen3.8-max` with a dated snapshot `qwen3.8-max-0902`. It reached general availability on **August 3, 2026**, having been previewed on **July 19, 2026** at the World AI Conference in Shanghai as `qwen3.8-max-preview` — a preview that shipped a headline parameter count and nothing else: no model card, no active-parameter figure, no scores.

Two things about it are worth more than any benchmark row.

The first is scale and sparsity. Qwen3.8-Max is a **2.4 trillion parameter [Mixture-of-Experts](https://howaiworks.ai/glossary/mixture-of-experts) model that activates 95 billion parameters per token**. Alibaba describes it as built on the Qwen3.5 architectural foundation and scaled up. Where Qwen3.7-Max published no architecture at all, this generation publishes the whole config, because you can download it.

The second is the licence, and it inverts what this page said three months ago. **The Qwen Max line is no longer closed.** The base checkpoint `Qwen/Qwen3.8-2.4T-A95B` went up on [huggingface.co/Qwen](https://huggingface.co/Qwen) on **August 12, 2026**, with an FP8 build alongside it, under a custom **Qwen3.8-Max License**. Qwen3.7-Max had no weights anywhere and Alibaba's flagship had looked like a permanent API product; that is no longer the description.

Read the licence before you plan around it, though. It is not Apache 2.0 — see [Limitations](#limitations) — and the downloadable checkpoint is not the model the API serves. Alibaba's own model card is explicit: "Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support." The open weights are **text-only and thinking-required**, with a 262,144-token native window. The hosted model adds vision and video input, a non-thinking path, and the full 1M window.

The third change from Qwen3.7-Max is modality. The previous flagship's documentation answered "Can I upload images or documents as input?" with "These models accept text only," and routed visual work to Qwen3-VL. Model Studio now lists `qwen3.8-max` under **image and video understanding** as well as text generation.

## Capabilities

- **Native vision and video input**: The hosted model takes text, image and video. Alibaba's launch material frames the 1M window in terms of what it holds — "hundred-page documents, full television series, or 100-hour livestreams" — rather than in tokens. This is new to the Max line as of 3.8.
- **Sparse MoE at 2.4T**: 95B active parameters per token out of 2.4T total. Serving cost tracks the active count, which is why a 2.4T model prices below the previous 3.7 generation rather than above it.
- **Selectable reasoning depth**: `xhigh` (the default) for tasks demanding thorough analysis, `medium` to balance accuracy against speed, `low` for cost- and latency-sensitive work. `preserve_thinking` is enabled by default across workloads.
- **1M token context**: The hosted API serves the full million; output is capped at 131,072 tokens.
- **Context caching**: Cached input reads at 10% of the standard input rate on both the Singapore and Beijing cards.
- **Long-horizon autonomy**: Alibaba's launch post reports a single unattended run of roughly 125 hours in which the model wrote about 7,600 lines of code, took over 1,100 actions and ran 33 rounds of GPU training. It is a vendor anecdote, not a benchmark, but it is the capability the model is being sold on.
- **Harness portability**: Alibaba claims comparable results across QwenWork, Claude Code, Codex, OpenClaw and Hermes — a claim about scaffold-independence that is straightforward to test on your own harness.
- **Downloadable base weights**: `Qwen3.8-2.4T-A95B` and an FP8 build, for teams that need to self-host and can meet the licence terms and the hardware bill.

## Technical Specifications

- **Hosted model IDs**: `qwen3.8-max` (rolling alias), `qwen3.8-max-0902` (snapshot, September 2, 2026), `qwen3.8-max-preview` (July 19, 2026)
- **Total parameters**: 2.4 trillion
- **Active parameters**: 95 billion per token
- **Architecture**: sparse [MoE](https://howaiworks.ai/glossary/mixture-of-experts); the open checkpoint's config declares 512 experts with 11 activated per layer (10 routed + 1 shared), in a hybrid stack alternating Gated DeltaNet blocks with gated attention
- **Context window**: 1,000,000 tokens hosted; 262,144 native on the open checkpoint, extensible to 1,010,000
- **Max output**: 131,072 tokens
- **Modalities**: text, image and video in, text out (hosted); text only on the open weights
- **Reasoning**: `xhigh` / `medium` / `low`; `preserve_thinking` on by default
- **Open weights**: `Qwen/Qwen3.8-2.4T-A95B` and `Qwen/Qwen3.8-2.4T-A95B-FP8`, published August 12, 2026
- **Licence**: Qwen3.8-Max License — free for most commercial use, with revenue- and MAU-gated conditions
- **Platforms**: Alibaba Cloud Model Studio and DashScope; QwenWork; Hugging Face and ModelScope for the weights

Alibaba has still not published a training-token count or a knowledge cutoff for Qwen3.8-Max. Those fields are absent from this page rather than estimated.

### The open-weight line in this generation

| Repository | Published | Type | Context | Licence |
|---|---|---|---|---|
| Qwen3.8-2.4T-A95B | 2026-08-12 | MoE, 2.4T total / 95B active, text-only | 262,144 (→1,010,000) | Qwen3.8-Max License |
| Qwen3.8-27B | 2026-08-14 | Dense 27B, vision-language | 262,144 (→1M via YaRN) | Apache 2.0 |

The smaller model kept Apache 2.0; the flagship did not. If permissive licensing is the hard requirement rather than open weights as such, the 27B is the model to build on. Qwen3.8-Flash-Next, a 125B MoE with 6B active, followed on August 26 under `qwen-community-1.0` — we covered its local-inference story in [Qwen3.8-Flash-Next Runs Locally in 75 GB of RAM](https://howaiworks.ai/blog/alibaba-qwen-3-8-flash-next-local-gguf).

### The prior release: Qwen3.7-Max

Qwen3.7-Max, launched May 20, 2026 at the Alibaba Cloud Summit in Hangzhou, remains available on Model Studio as `qwen3.7-max` with the `qwen3.7-max-2026-05-20` and `qwen3.7-max-2026-06-08` snapshots. It was **closed-weight and API-only**, text-only, with hybrid thinking on by default and a 1M token [context window](https://howaiworks.ai/glossary/context-window). Its Singapore rate card is $2.50 input and $7.50 output per 1M tokens — above the 3.8 flagship on both sides. The lower-cost `qwen3.7-plus` sibling is still the tiered option: $0.40 / $1.60 per 1M up to 256K input tokens, rising to $1.20 / $4.80 beyond it. There is no `qwen3.8-plus`; the cheap slot in this generation is `qwen3.8-flash`.

## Use Cases

- **Long-horizon coding agents**: Multi-day unattended runs with tool use and self-correction are the workload Alibaba built and benchmarked for. The 1M window plus `xhigh` reasoning is the intended configuration.
- **Video and document understanding at length**: Native video input against a million-token window suits transcript-plus-frames analysis of long recordings — meetings, broadcasts, livestream archives — without a separate vision model in the pipeline.
- **Whole-repository reasoning**: A 1M window holds a substantial codebase in a single request, removing much of the retrieval scaffolding.
- **Sovereign or air-gapped deployment**: Newly possible on this line. `Qwen3.8-2.4T-A95B` can be self-hosted under vLLM, SGLang or comparable runtimes — subject to the licence gates and to finding the cluster.
- **Cost-sensitive frontier work**: At $2.00 / $6.00 it undercuts the Western frontier tier substantially while claiming comparable scores. Whether the scores hold on your tasks is the whole question.
- **Multilingual enterprise workloads**: The Qwen series' long-standing strength in Chinese and English, now with a self-hosted fallback.

## Performance / Benchmarks

Alibaba published a first-party evaluation table with the launch post, ["Qwen3.8-Max: A New Bar for Coding and Cowork"](https://www.alibabacloud.com/blog/qwen3-8-max-a-new-bar-for-coding-and-cowork_603421) (Alibaba Cloud Community, August 3, 2026). **All figures below are Alibaba-reported**, measured by Alibaba, at launch.

A methodological note before the numbers: Alibaba renders its benchmark tables as images, and the [qwen.ai](https://qwen.ai/) blog serves no static text at all. The scores below are transcribed from the launch material and cross-checked against independent write-ups of it; treat the second decimal with more suspicion than you would a machine-readable table. Comparator scores are reproduced as Alibaba stated them, under the names Alibaba used.

| Benchmark | Qwen3.8-Max | Comparators cited by Alibaba |
|---|---|---|
| Terminal Bench 2.1 | 86.6 | GPT-5.6 Sol 88.8; Claude Opus 4.8 84.6; Fable 5 84.6 |
| SWE-bench Pro | 67.7 | Fable 5 80.0; Opus 4.8 69.2; GPT-5.6 Sol 64.6 |
| PaperBench | 93.0 | GPT-5.6 Sol 90.5; Fable 5 88.8; Opus 4.8 80.3 |
| GPQA Diamond | 92.6 | GPT-5.6 Sol 94.1; Fable 5 92.6; Opus 4.8 92.0 |
| IFBench | 82.8 | GPT-5.6 Sol 72.7; Fable 5 63.5; Opus 4.8 62.2 |
| HLE | 43.6 | Fable 5 53.3; GPT-5.6 Sol 47.2; Opus 4.8 45.7 |
| OSWorld-Verified | 86.1 | Fable 5 85.0; GPT-5.6 Sol Max 83.2; Gemini 3.1 Pro 76.2 |
| MathVision | 95.2 | — |
| LogicVista | 91.9 | — |

The shape of that table is more informative than any single row. Qwen3.8-Max leads on instruction-following (IFBench, by twenty points over the Claude entries), on research reproduction (PaperBench) and on computer use (OSWorld-Verified), and trails on the two hardest reasoning and software-engineering rows — HLE and SWE-bench Pro — by wide margins. It is not a uniform frontier claim and Alibaba does not present it as one.

Four caveats travel with these numbers:

- **They are vendor-run and unreplicated.** No independent evaluation of Qwen3.8-Max had been published at the time of writing.
- **Most coding rows used the Claude Code harness**, not a neutral scaffold. Harness choice moves agentic coding scores materially.
- **Several benchmarks are Alibaba's own constructions** and cannot be audited by a third party — including the E-Commerce Bench simulation Alibaba highlights, where it reports a ¥416,252 closing balance, a 4.16x return, and a 38% margin over second-place GLM-5.2.
- **The comparator names are Alibaba's.** Where Alibaba cites "Claude Opus 4.8," "Fable 5" or "GPT-5.6 Sol," the configuration behind that label is not independently documented in the post.

For the September snapshot, Alibaba reports `qwen3.8-max-0902` at 1,691 on Code Arena WebDev, and large jumps on two coding evaluations — TerminalBench 3.0 from 11.3 to 29.0 and ProgramBench from 10.5 to 28.0. Those are first-party figures on a post-training refresh, and the low absolute values are the interesting part: both benchmarks are hard enough that the frontier is still under 30.

## Limitations

- **Open weight is not Apache 2.0.** The Qwen3.8-Max License is broadly permissive but carries two gates: display the model name in your UI above 100M MAU or US$20M monthly revenue, and negotiate a separate licence if you run a Model-as-a-Service or "AI Work Assistant" business above US$50M revenue in any twelve months. Legal review is not optional for anyone near those thresholds.
- **The open checkpoint is a different model from the API.** No vision, no video, no non-thinking mode, and 262,144 native context rather than 1M. Benchmarks measured on the hosted model do not transfer to a self-hosted deployment.
- **Self-hosting 2.4T is a datacenter problem.** Sparse activation cuts compute per token, not storage. The full-precision weights are a multi-node deployment; FP8 helps but does not make this a single-server model.
- **No published knowledge cutoff and no training-token count.** Ground time-sensitive queries with retrieval or tool use.
- **Benchmarks are vendor-run, harness-favourable and partly internal.** See the four caveats above.
- **Thinking is on by default at the deepest setting.** `xhigh` is the default reasoning effort, so latency and output-token spend both run high unless you explicitly step down to `medium` or `low`.
- **Regional price divergence, and two currencies.** Singapore is quoted natively in USD, Beijing natively in CNY. Cost models must account for which region serves the request and for the exchange rate Alibaba applies when restating one card in the other's currency.
- **Preview-era documentation was thin, and some of it persists.** The July preview shipped without a model card; third-party summaries written in that window still circulate with no active-parameter figure and no scores.

## Pricing & Access

Alibaba quotes each region in its own currency, and the two rate cards are not conversions of one another. **The International (Singapore) card is natively in US dollars. The Chinese mainland (North China 2, Beijing) card is natively in Chinese yuan.** Each pricing page restates the other region's figures in its own currency, which is where the long decimals come from. Prices are per 1M tokens.

### International (Singapore) — quoted in USD

Source: [Alibaba Cloud Model Studio pricing](https://www.alibabacloud.com/help/en/model-studio/model-pricing).

| Model | Input | Cached input | Output |
|---|---|---|---|
| `qwen3.8-max` | $2.00 | $0.20 | $6.00 |
| `qwen3.8-flash` | $0.15 | — | $0.47 |
| `qwen3.7-max` (prior flagship) | $2.50 | $0.25 | $7.50 |

`qwen3.8-max` is flat across its entire 1M window — no tiering. Cached input reads at 10% of the standard rate.

### Chinese mainland — North China 2 (Beijing) — quoted in CNY

Source: [阿里云百炼模型计费](https://help.aliyun.com/zh/model-studio/model-pricing) ("Alibaba Cloud Model Studio model pricing"), section 华北2（北京）.

| Model | Input | Output |
|---|---|---|
| `qwen3.8-max` | ¥12 | ¥36 |
| `qwen3.8-flash` | ¥0.8 | ¥2.7 |

Alibaba's English page restates the mainland `qwen3.8-max` row as $1.65 / $4.951. Those are converted figures, not the vendor's native quote — the round yuan numbers are.

**Access options:**

- **Alibaba Cloud Model Studio** — the primary managed API, and where the snapshots are pinned
- **DashScope** — Alibaba's model-serving platform
- **QwenWork** — Alibaba's workplace agent platform, in public beta since the August 3 launch
- **Hugging Face and ModelScope** — `Qwen3.8-2.4T-A95B` and its FP8 build, under the Qwen3.8-Max License
- **Third-party aggregators** — resell the hosted API

## Ecosystem & Tools

- **[Qwen3.8-Max: A New Bar for Coding and Cowork](https://www.alibabacloud.com/blog/qwen3-8-max-a-new-bar-for-coding-and-cowork_603421)** — Alibaba's launch post and the primary source for the benchmark table
- **[Alibaba unveils Qwen3.8-Max](https://www.alibabacloud.com/en/press-room/alibaba-unveils-qwen3-8-max)** — the press-room announcement, August 3, 2026
- **[Alibaba Cloud Model Studio](https://www.alibabacloud.com/help/en/model-studio/models)** — supported models, capabilities and snapshot IDs
- **[Model Studio pricing (English)](https://www.alibabacloud.com/help/en/model-studio/model-pricing)** — per-region rate cards
- **[百炼模型计费 (Chinese)](https://help.aliyun.com/zh/model-studio/model-pricing)** — the mainland rate card as Alibaba natively quotes it, in CNY
- **[Qwen/Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B)** — the open base checkpoint, its config, and the licence text
- **[Qwen on Hugging Face](https://huggingface.co/Qwen)** — the full open-weight catalogue, including Qwen3.8-27B under Apache 2.0
- **[QwenLM on GitHub](https://github.com/QwenLM)** — inference recipes and model repositories
- **[ModelScope](https://modelscope.cn/organization/qwen)** — Alibaba's model community

## Community & Resources

- [Supported models and capabilities overview](https://www.alibabacloud.com/help/en/model-studio/models)
- [Text generation models on Model Studio](https://www.alibabacloud.com/help/en/model-studio/text-generation-model/)
- [Qwen](https://qwen.ai/)
- Our write-up of the smaller sibling: [Qwen3.8-Flash-Next Runs Locally in 75 GB of RAM](https://howaiworks.ai/blog/alibaba-qwen-3-8-flash-next-local-gguf)
- Compare with [DeepSeek V4](https://howaiworks.ai/models/deepseek), [Kimi K3](https://howaiworks.ai/models/kimi), [GLM-5.3](https://howaiworks.ai/models/glm), and [Tencent Hy3](https://howaiworks.ai/models/hunyuan)

## Frequently Asked Questions

### When was Qwen3.8-Max released?

August 3, 2026, when Alibaba moved it to general availability on Model Studio and published its benchmark table. A preview, qwen3.8-max-preview, had been shown on July 19, 2026 at the World AI Conference in Shanghai with the total parameter count but no active-parameter figure, no model card and no scores.

### Is Qwen3.8-Max open source?

It is open weight, not open source, and the distinction matters. The base checkpoint Qwen3.8-2.4T-A95B was published on huggingface.co/Qwen on August 12, 2026 under a custom Qwen3.8-Max License — not Apache 2.0. This reverses the posture of Qwen3.7-Max, which was closed and API-only.

### What does the Qwen3.8-Max License restrict?

Two revenue-gated carve-outs sit on top of otherwise broad commercial rights. A product with more than 100,000,000 monthly active users or more than US$20,000,000 in monthly revenue must display the model name prominently in its interface. A Model-as-a-Service or AI Work Assistant business whose revenue exceeds US$50,000,000 in any consecutive twelve months needs a separate licence from Alibaba before commercial use. Purely internal use that exposes nothing to third parties is excluded from the MaaS trigger.

### Are the open weights the same model as the hosted Qwen3.8-Max?

No, and Qwen says so on the model card: Qwen3.8-Max is "the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support." The open checkpoint is text-only, requires thinking mode, and ships with a 262,144-token native window extensible to 1,010,000. The hosted model adds vision and video input, a non-thinking path, and the full 1M window.

### Does Qwen3.8-Max accept images and video?

The hosted model does. Alibaba Cloud Model Studio lists qwen3.8-max under both text generation and image and video understanding, and Alibaba's launch material describes it processing hundred-page documents, full television series and 100-hour livestreams. This is a change from Qwen3.7-Max, whose documentation stated flatly that the Max snapshots accepted text only.

### How much does Qwen3.8-Max cost?

In the International (Singapore) region, which Alibaba quotes natively in US dollars, it is $2.00 per 1M input tokens and $6.00 per 1M output tokens, with cached input at $0.20 per 1M — a 90% discount. That is a 20% cut on both sides against Qwen3.7-Max's $2.50 / $7.50. The Chinese mainland (Beijing) card is quoted natively in yuan at ¥12 and ¥36 per 1M tokens.

### What is the context window for Qwen3.8-Max?

1 million tokens on the hosted API, with output capped at 131,072 tokens. The downloadable Qwen3.8-2.4T-A95B checkpoint is natively 262,144 tokens and extensible to 1,010,000, so a self-hosted deployment does not get the full window for free.

### What is qwen3.8-max-0902?

A dated snapshot published September 2, 2026. Alibaba describes it as a post-training refresh on the same architecture, aimed at engineering-scale coding, multi-tool orchestration in agent runs, and chart and document parsing. It does not replace the rolling qwen3.8-max alias; both remain available, and pinning the snapshot is what gives reproducible behaviour.

### How many parameters does Qwen3.8-Max have?

2.4 trillion in total with 95 billion activated per token — a sparse Mixture-of-Experts design. The open checkpoint's config shows 512 experts with 11 activated per layer, 10 routed plus 1 shared, in a hybrid stack that alternates Gated DeltaNet blocks with gated attention.

### Are the Qwen3.8-Max benchmark scores independent?

No. Every figure Alibaba published on August 3, 2026 is vendor-run and unreplicated, several of the benchmarks are Alibaba's own constructions, and most coding rows were measured inside the Claude Code harness rather than a neutral scaffold. Treat them as an upper bound and re-run the evaluations that matter to you.

## Related

### Related models

- [GPT-6 Astra](https://howaiworks.ai/models/gpt)
- [Claude Fable 5.1](https://howaiworks.ai/models/claude-fable)
- [Gemini 3.8 Flash](https://howaiworks.ai/models/gemini)
- [Grok 4.6](https://howaiworks.ai/models/grok)
- [DeepSeek V4](https://howaiworks.ai/models/deepseek)
- [Kimi K3](https://howaiworks.ai/models/kimi)

---

Source: https://howaiworks.ai/models/qwen — HowAIWorks.ai
