ApX logoApX logo

Kimi K3

Active Parameters

2.8T

Context Length

1.05M

Modality

Multimodal

Architecture

Mixture of Experts (MoE)

License

Open Weights

Release Date

27 Jul 2026

Knowledge Cutoff

Jun 2024

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

5886.41 GB VRAM

Consumer

526x RTX 4090

24GB VRAM

Datacenter

113x NVIDIA A100

80GB VRAM

Apple Silicon

122x Apple M3 Max

128GB VRAM

1,048,576 tokens

10914.34 GB VRAM

Consumer

1259x RTX 4090

24GB VRAM

Datacenter

243x NVIDIA A100

80GB VRAM

Apple Silicon

309x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 7.2k · Context: 1.05M · Vocab: 163.8kx 93 layersRMSNormPre-AttentionMulti-Layer Attention96Q / 96KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (16/896 experts)SwiGLUIntermediate: 3.1k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Kimi K3 available.

Rankings

Overall Rank

-

Coding Rank

-

About Kimi K3

Kimi K3 is a flagship 2.8-trillion-parameter sparse Mixture-of-Experts (MoE) model developed by Moonshot AI, engineered for long-horizon agentic workflows and large-scale multimodal reasoning. The architecture utilizes a massive scale of 896 total experts, of which 16 are active per token, resulting in approximately 104 billion active parameters during inference. A defining feature of Kimi K3 is its 1-million-token context window, supported by a novel hybrid attention mechanism that combines Kimi Delta Attention (KDA) for linear-time sequence processing with Multi-Head Latent Attention (MLA) layers to maintain high-fidelity retrieval across extensive sequences.

The model introduces several architectural innovations to ensure stability and efficiency at the 3-trillion-parameter scale. This includes Attention Residuals (AttnRes), which allow layers to selectively sample representations from earlier depths rather than relying on uniform accumulation, and the Stable LatentMoE framework, which utilizes Quantile Balancing to prevent expert starvation without traditional auxiliary losses. Kimi K3 is natively multimodal, featuring a custom-trained vision encoder that allows the model to process text, images, and video within a single unified representation space, making it particularly effective for vision-in-the-loop tasks like GUI automation and complex debugging.

Optimized for high-concurrency environments, Kimi K3 supports advanced inference techniques including quantization-aware training from the supervised fine-tuning stage. It is designed to be served in large-scale cluster configurations, typically requiring 64 or more accelerators for optimal performance. The model's primary use cases center on autonomous software engineering, multi-document synthesis, and complex research orchestration where maintaining coherence over hours or days of execution is required.

Technical Specifications

Attention

Attention Structure

Multi-Layer Attention

Attention Heads

96

Key-Value Heads

96

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

Yes

Linear Attention Ratio

74.2%

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

7,168

Number of Layers

93

FFN Intermediate Size (Dense)

33,792

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

163,840

Mixture of Experts

Total Expert Parameters

104.0B

Number of Experts

896

Active Experts

16

Shared Experts

2

FFN Intermediate Size (per Expert)

3,072

Dense Layers Before MoE

1

About Kimi K3

Kimi K3 is a flagship Mixture-of-Experts (MoE) multimodal model series developed by Moonshot AI.


Other Kimi K3 Models
  • No related models available
Kimi K3: Specifications and GPU VRAM Requirements