Active Parameters
2.4T
Context Length
262K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
Apache 2.0
Release Date
12 Aug 2026
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
427x RTX 4090
24GB VRAM
Datacenter
94x NVIDIA A100
80GB VRAM
Apple Silicon
98x Apple M3 Max
128GB VRAM
262,144 tokens
Consumer
439x RTX 4090
24GB VRAM
Datacenter
96x NVIDIA A100
80GB VRAM
Apple Silicon
101x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Qwen3.8 2.4T A95B available.
Overall Rank
-
Coding Rank
-
Qwen3.8-2.4T-A95B is Alibaba's flagship open-weight Mixture-of-Experts (MoE) foundation model with 2.4 trillion total parameters and 95 billion active parameters per token (512 routed experts with 10 active + 1 shared expert). Built on a 92-layer hybrid backbone interleaving Gated DeltaNet linear attention with standard Gated Attention (full attention at every 4th layer), it delivers frontier reasoning, coding, and autonomous agent capabilities with native 262K context (extensible to 1M) and mandatory thinking mode.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
64
Key-Value Heads
4
Attention Head Dimension
256
Position Embedding
ROPE
RoPE Theta
10,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
75.0%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
8,192
Number of Layers
92
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
248,320
Mixture of Experts
Total Expert Parameters
95.0B
Number of Experts
512
Active Experts
11
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
-
Alibaba's Qwen 3.8 generation represents the frontier hybrid Mixture-of-Experts architecture designed for coding, professional work, research, and long-horizon agentic tasks. It features a 2.4-trillion parameter architecture (95B active per token) combining Gated DeltaNet linear attention with standard Gated Attention, available both as open weights and as a hosted flagship service.
APX AI
Online