Active Parameters
2.8T
Context Length
1.05M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Open Weights
Release Date
27 Jul 2026
Knowledge Cutoff
Jun 2024
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
526x RTX 4090
24GB VRAM
Datacenter
113x NVIDIA A100
80GB VRAM
Apple Silicon
122x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
1259x RTX 4090
24GB VRAM
Datacenter
243x NVIDIA A100
80GB VRAM
Apple Silicon
309x Apple M3 Max
128GB VRAM
Rank
#12
| Benchmark | Score | Rank |
|---|---|---|
Reasoning | 0.91 | ⭐ 4 |
Web Development max | 1674 | ⭐ 5 |
Agentic Coding | 0.62 | ⭐ 6 |
General | 0.79 | ⭐ 7 |
Agent Arena max | 0.09 | 8 |
LiveBench Average | 0.79 | ⭐ 8 |
Data Analysis | 0.79 | 12 |
General Text max | 1489 | ⭐ 13 |
Coding | 0.81 | 13 |
Mathematics | 0.84 | 55 |
Overall Rank
#12
Coding Rank
#19
Kimi K3 is Moonshot AI's flagship 2.8T-parameter Mixture-of-Experts multimodal model engineered for long-horizon agentic workflows and complex reasoning. Featuring a 1M token context window powered by hybrid linear attention, it excels at autonomous coding and vision-in-the-loop tasks.
Attention
Attention Structure
Multi-Layer Attention
Attention Heads
96
Key-Value Heads
96
Attention Head Dimension
128
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
74.2%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
7,168
Number of Layers
93
FFN Intermediate Size (Dense)
33,792
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
163,840
Mixture of Experts
Total Expert Parameters
104.0B
Number of Experts
896
Active Experts
16
Shared Experts
2
FFN Intermediate Size (per Expert)
3,072
Dense Layers Before MoE
1
Kimi K3 is a flagship Mixture-of-Experts (MoE) multimodal model series developed by Moonshot AI.
APX AI
Online