Active Parameters
16B
Context Length
128K
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
MIT
Release Date
10 Apr 2025
Knowledge Cutoff
-
API Pricing (per 1M)
Self-hosted only
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
2x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
3x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Kimi-VL-A3B-Instruct available.
Overall Rank
-
Coding Rank
-
Kimi-VL-A3B-Instruct is a vision-language Mixture-of-Experts model by Moonshot AI activating 2.8B parameters for high-resolution perceptual tasks. It excels at document parsing, GUI agent interactions, and video comprehension over a 128K context window.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
16
Key-Value Heads
16
Attention Head Dimension
-
Position Embedding
Absolute Position Embedding
RoPE Theta
800,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
2,048
Number of Layers
24
FFN Intermediate Size (Dense)
1,408
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
163,840
Mixture of Experts
Total Expert Parameters
3.0B
Number of Experts
384
Active Experts
8
Shared Experts
2
FFN Intermediate Size (per Expert)
1,408
Dense Layers Before MoE
1
Kimi-VL by Moonshot AI is an efficient, open-source Mixture-of-Experts vision-language model. It employs a native-resolution MoonViT encoder and an MoE language model, activating 2.8 billion parameters. The model handles high-resolution visual inputs and processes contexts up to 128K tokens. A "Thinking" variant provides enhanced long-horizon reasoning.
APX AI
Online