Active Parameters
2.4T
Context Length
1M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Proprietary
Release Date
2 Aug 2026
Knowledge Cutoff
-
No evaluation benchmarks for Qwen3.8 Max available.
Overall Rank
-
Coding Rank
-
Qwen3.8-Max is Alibaba Cloud's flagship hosted foundation model based on the 2.4T parameter (95B active) Mixture-of-Experts architecture. Featuring a hybrid Gated DeltaNet linear attention and standard Gated Attention design, it provides native vision input capabilities, default 1M token context window, non-thinking and thinking mode toggles, and built-in agent tooling for complex software engineering, research reproduction, and long-horizon autonomous tasks.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
64
Key-Value Heads
4
Attention Head Dimension
256
Position Embedding
ROPE
RoPE Theta
10,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
75.0%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
8,192
Number of Layers
92
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
248,320
Mixture of Experts
Total Expert Parameters
95.0B
Number of Experts
512
Active Experts
11
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
-
Alibaba's Qwen 3.8 generation represents the frontier hybrid Mixture-of-Experts architecture designed for coding, professional work, research, and long-horizon agentic tasks. It features a 2.4-trillion parameter architecture (95B active per token) combining Gated DeltaNet linear attention with standard Gated Attention, available both as open weights and as a hosted flagship service.
APX AI
Online