Active Parameters
2.4T (Estimated)
Context Length
1M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Proprietary
Release Date
2 Aug 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $2.00 · Output: $6.00
Rank
#16
| Benchmark | Score | Rank |
|---|---|---|
Agentic Index | 0.56 | 🥉 3 |
Web Development | 1671 | 4 |
Agentic Coding | 0.65 | 6 |
Graduate-Level QA | 0.926 | 11 |
General | 0.78 | 13 |
LiveBench Average | 0.78 | 13 |
Data Analysis | 0.78 | 17 |
Coding Index | 0.76 | 17 |
Reasoning | 0.88 | 18 |
General Text | 1481 | 20 |
Agent Arena | 0.03 | 21 |
Intelligence Index | 0.45 | 21 |
Mathematics | 0.91 | 23 |
Coding | 0.73 | 43 |
Overall Rank
#16
Coding Rank
#35
Qwen3.8-Max is Alibaba Cloud's flagship hosted foundation model based on the 2.4T parameter (95B active) Mixture-of-Experts architecture. Featuring a hybrid Gated DeltaNet linear attention and standard Gated Attention design, it provides native vision input capabilities, default 1M token context window, non-thinking and thinking mode toggles, and built-in agent tooling for complex software engineering, research reproduction, and long-horizon autonomous tasks.
Architecture specifications are undisclosed for proprietary models.
Alibaba's Qwen 3.8 generation represents the frontier hybrid Mixture-of-Experts architecture designed for coding, professional work, research, and long-horizon agentic tasks. It features a 2.4-trillion parameter architecture (95B active per token) combining Gated DeltaNet linear attention with standard Gated Attention, available both as open weights and as a hosted flagship service.
Assistant
Online