Active Parameters
2.4T
Context Length
1.05M
Modality
Text
Architecture
Mixture of Experts (MoE)
License
other
Release Date
8 Aug 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $2.00 · Output: $6.00
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
427x RTX 4090
24GB VRAM
Datacenter
94x NVIDIA A100
80GB VRAM
Apple Silicon
98x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
475x RTX 4090
24GB VRAM
Datacenter
103x NVIDIA A100
80GB VRAM
Apple Silicon
110x Apple M3 Max
128GB VRAM
Rank
#13
| Benchmark | Score | Rank |
|---|---|---|
Agentic Index | 0.50 | 16 |
Intelligence Index | 0.40 | 30 |
Coding Index | 0.72 | 33 |
Overall Rank
#13
Coding Rank
#31
Qwen 3.8 2.4T A95B is a massive open-weight sparse mixture-of-experts model from Alibaba activating 95B parameters out of 2.4T total. It provides frontier-level general reasoning, coding, and multilingual performance as the open variant of Qwen 3.8 Max.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
64
Key-Value Heads
4
Attention Head Dimension
256
Position Embedding
ROPE
RoPE Theta
10,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
75.0%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
8,192
Number of Layers
92
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
248,320
Mixture of Experts
Total Expert Parameters
95.0B
Number of Experts
512
Active Experts
10
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
0
APX AI
Online