Active Parameters
2.4T (Estimated)
Context Length
1M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Proprietary
Release Date
2 Aug 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $2.00 · Output: $6.00
Rank
#12
| Benchmark | Score | Rank |
|---|---|---|
Web Development | 1689 | 🥈 2 |
Agentic Coding | 0.65 | 4 |
General | 0.78 | 9 |
LiveBench Average | 0.78 | 9 |
Graduate-Level QA | 0.926 | 10 |
Data Analysis | 0.78 | 13 |
Reasoning | 0.88 | 15 |
Mathematics | 0.91 | 17 |
Agent Arena | 0.05 | 18 |
Intelligence Index | 0.47 | 18 |
Agentic Index | 0.50 | 18 |
General Text | 1480 | 20 |
Coding Index | 0.72 | 31 |
Coding | 0.73 | 35 |
Overall Rank
#12
Coding Rank
#22
Qwen3.8-Max is Alibaba Cloud's flagship hosted foundation model based on the 2.4T parameter (95B active) Mixture-of-Experts architecture. Featuring a hybrid Gated DeltaNet linear attention and standard Gated Attention design, it provides native vision input capabilities, default 1M token context window, non-thinking and thinking mode toggles, and built-in agent tooling for complex software engineering, research reproduction, and long-horizon autonomous tasks.
Architecture specifications are undisclosed for proprietary models.
Alibaba's Qwen 3.8 generation represents the frontier hybrid Mixture-of-Experts architecture designed for coding, professional work, research, and long-horizon agentic tasks. It features a 2.4-trillion parameter architecture (95B active per token) combining Gated DeltaNet linear attention with standard Gated Attention, available both as open weights and as a hosted flagship service.
APX AI
Online