Active Parameters
2.8T
Context Length
1.05M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Open Weights
Release Date
27 Jul 2026
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
526x RTX 4090
24GB VRAM
Datacenter
113x NVIDIA A100
80GB VRAM
Apple Silicon
122x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
1259x RTX 4090
24GB VRAM
Datacenter
243x NVIDIA A100
80GB VRAM
Apple Silicon
309x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Kimi K3 available.
Overall Rank
-
Coding Rank
-
Kimi K3 is a flagship open-weights 2.8T parameter sparse Mixture-of-Experts (MoE) multimodal model developed by Moonshot AI. It features a 1M token context window, 896 total experts with 16 active experts per token, hybrid linear attention (Kimi Delta Attention), and native support for text and vision inputs.
Attention
Attention Structure
Multi-Layer Attention
Attention Heads
96
Key-Value Heads
96
Attention Head Dimension
128
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
74.2%
Normalization
RMS Normalization
Activation Function
-
Dimensions
Hidden Dimension Size
7,168
Number of Layers
93
FFN Intermediate Size (Dense)
33,792
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
163,840
Mixture of Experts
Total Expert Parameters
32.0B
Number of Experts
896
Active Experts
16
Shared Experts
2
FFN Intermediate Size (per Expert)
3,072
Dense Layers Before MoE
1
Kimi K3 is a flagship Mixture-of-Experts (MoE) multimodal model series developed by Moonshot AI.
APX AI
Online