ApX logoApX logo

MiMo V2.6 Pro RL

Active Parameters

1.02T

Context Length

1.05M

Modality

Text

Architecture

Mixture of Experts (MoE)

License

MIT License

Release Date

21 Sept 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Self-hosted only

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

2143.96 GB VRAM

Consumer

143x RTX 4090

24GB VRAM

Datacenter

35x NVIDIA A100

80GB VRAM

Apple Silicon

31x Apple M3 Max

128GB VRAM

1,048,576 tokens

2617.02 GB VRAM

Consumer

183x RTX 4090

24GB VRAM

Datacenter

44x NVIDIA A100

80GB VRAM

Apple Silicon

40x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 6.1k · Context: 1.05M · Vocab: 152.6kx 70 layersRMSNormPre-AttentionSliding-Window Attention128Q / 8KV heads · SW: 128Head dim: 192+RMSNormPre-FFNSparse MoE FFN (8/384 experts)SwiGLUIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for MiMo V2.6 Pro RL available.

Rankings

Overall Rank

-

Coding Rank

-

About MiMo V2.6 Pro RL

MiMo V2.6 Pro RL is Xiaomi's flagship reasoning-oriented model trained using reinforcement learning for complex problem solving and agent workflows. It achieves enhanced performance across coding, mathematics, and complex multi-step tasks.

Technical Specifications

Attention

Attention Structure

Single-Head Attention

Attention Heads

128

Key-Value Heads

8

Attention Head Dimension

192

Position Embedding

ROPE

RoPE Theta

10,000,000

Sliding Window Attention

Yes

Sliding Window Size

128

Sliding Window Ratio

85.7%

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

6,144

Number of Layers

70

FFN Intermediate Size (Dense)

16,384

Multi-Token Prediction Heads

5

Tokenizer

Vocabulary Size

152,576

Mixture of Experts

Total Expert Parameters

42.0B

Number of Experts

384

Active Experts

8

Shared Experts

-

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

1

About MiMo V2

MiMo-V2-Flash is a Mixture-of-Experts (MoE) model with hybrid attention architecture designed for high-speed reasoning and agentic workflows. It features Multi-Token Prediction (MTP) to achieve state-of-the-art performance while significantly reducing inference costs. The model is optimized for long-context modeling and efficient inference.


Other MiMo V2 Models
MiMo V2.6 Pro RL: Specifications and GPU VRAM Requirements