Active Parameters
15B
Context Length
256K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
MIT
Release Date
10 Dec 2025
Knowledge Cutoff
Dec 2024
API Pricing (per 1M)
Input: $0.10 · Output: $0.30
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
2x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
256,000 tokens
Consumer
4x RTX 4090
24GB VRAM
Datacenter
2x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
Rank
#71
| Benchmark | Score | Rank |
|---|---|---|
Professional Knowledge | 0.849 | 15 |
Graduate-Level QA | 0.837 | 48 |
Web Development | 1331 | 83 |
Coding Index | 0.50 | 84 |
General Text | auto 1386 Standard 1392 | 113 105 |
Intelligence Index | 0.26 | 109 |
Overall Rank
#71
Coding Rank
#79
MiMo V2 Flash is a 309B Mixture-of-Experts model by Xiaomi activating 15B parameters for high-speed reasoning and software engineering. Powered by hybrid sliding-window attention and Multi-Token Prediction, it supports contexts up to 256K tokens.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
64
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
Absolute Position Embedding
RoPE Theta
640,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
4,096
Number of Layers
48
FFN Intermediate Size (Dense)
11,008
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
151,680
Mixture of Experts
Total Expert Parameters
309.0B
Number of Experts
256
Active Experts
8
Shared Experts
-
FFN Intermediate Size (per Expert)
-
Dense Layers Before MoE
-
MiMo-V2-Flash is a Mixture-of-Experts (MoE) model with hybrid attention architecture designed for high-speed reasoning and agentic workflows. It features Multi-Token Prediction (MTP) to achieve state-of-the-art performance while significantly reducing inference costs. The model is optimized for long-context modeling and efficient inference.
APX AI
Online