Active Parameters
2.8T
Context Length
1.05M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Open Weights
Release Date
27 Jul 2026
Knowledge Cutoff
Jun 2024
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
526x RTX 4090
24GB VRAM
Datacenter
113x NVIDIA A100
80GB VRAM
Apple Silicon
122x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
1259x RTX 4090
24GB VRAM
Datacenter
243x NVIDIA A100
80GB VRAM
Apple Silicon
309x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Kimi K3 available.
Overall Rank
-
Coding Rank
-
Kimi K3 is a flagship 2.8-trillion-parameter sparse Mixture-of-Experts (MoE) model developed by Moonshot AI, engineered for long-horizon agentic workflows and large-scale multimodal reasoning. The architecture utilizes a massive scale of 896 total experts, of which 16 are active per token, resulting in approximately 104 billion active parameters during inference. A defining feature of Kimi K3 is its 1-million-token context window, supported by a novel hybrid attention mechanism that combines Kimi Delta Attention (KDA) for linear-time sequence processing with Multi-Head Latent Attention (MLA) layers to maintain high-fidelity retrieval across extensive sequences.
The model introduces several architectural innovations to ensure stability and efficiency at the 3-trillion-parameter scale. This includes Attention Residuals (AttnRes), which allow layers to selectively sample representations from earlier depths rather than relying on uniform accumulation, and the Stable LatentMoE framework, which utilizes Quantile Balancing to prevent expert starvation without traditional auxiliary losses. Kimi K3 is natively multimodal, featuring a custom-trained vision encoder that allows the model to process text, images, and video within a single unified representation space, making it particularly effective for vision-in-the-loop tasks like GUI automation and complex debugging.
Optimized for high-concurrency environments, Kimi K3 supports advanced inference techniques including quantization-aware training from the supervised fine-tuning stage. It is designed to be served in large-scale cluster configurations, typically requiring 64 or more accelerators for optimal performance. The model's primary use cases center on autonomous software engineering, multi-document synthesis, and complex research orchestration where maintaining coherence over hours or days of execution is required.
Attention
Attention Structure
Multi-Layer Attention
Attention Heads
96
Key-Value Heads
96
Attention Head Dimension
128
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
74.2%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
7,168
Number of Layers
93
FFN Intermediate Size (Dense)
33,792
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
163,840
Mixture of Experts
Total Expert Parameters
104.0B
Number of Experts
896
Active Experts
16
Shared Experts
2
FFN Intermediate Size (per Expert)
3,072
Dense Layers Before MoE
1
Kimi K3 is a flagship Mixture-of-Experts (MoE) multimodal model series developed by Moonshot AI.
APX AI
Online