ApX logoApX logo

Kimi-VL-A3B-Instruct

Active Parameters

16B

Context Length

128K

Modality

Multimodal

Architecture

Mixture of Experts (MoE)

License

MIT

Release Date

10 Apr 2025

Knowledge Cutoff

-

API Pricing (per 1M)

Self-hosted only

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

35.31 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

128,000 tokens

61.52 GB VRAM

Consumer

3x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 2k · Context: 128K · Vocab: 163.8kx 24 layersRMSNormPre-AttentionMulti-Head Attention16Q / 16KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (8/384 experts)SwiGLUIntermediate: 1.4k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Kimi-VL-A3B-Instruct available.

Rankings

Overall Rank

-

Coding Rank

-

About Kimi-VL-A3B-Instruct

Kimi-VL-A3B-Instruct is a vision-language Mixture-of-Experts model by Moonshot AI activating 2.8B parameters for high-resolution perceptual tasks. It excels at document parsing, GUI agent interactions, and video comprehension over a 128K context window.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

16

Key-Value Heads

16

Attention Head Dimension

-

Position Embedding

Absolute Position Embedding

RoPE Theta

800,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

2,048

Number of Layers

24

FFN Intermediate Size (Dense)

1,408

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

163,840

Mixture of Experts

Total Expert Parameters

3.0B

Number of Experts

384

Active Experts

8

Shared Experts

2

FFN Intermediate Size (per Expert)

1,408

Dense Layers Before MoE

1

About Kimi-VL

Kimi-VL by Moonshot AI is an efficient, open-source Mixture-of-Experts vision-language model. It employs a native-resolution MoonViT encoder and an MoE language model, activating 2.8 billion parameters. The model handles high-resolution visual inputs and processes contexts up to 128K tokens. A "Thinking" variant provides enhanced long-horizon reasoning.


Other Kimi-VL Models
  • No related models available