ApX logoApX logo

K2 Horizon MoVA 36B A4B

Active Parameters

36B

Context Length

524K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

1 Sept 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Self-hosted only

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

77.31 GB VRAM

Consumer

4x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

524,288 tokens

185.33 GB VRAM

Consumer

9x RTX 4090

24GB VRAM

Datacenter

3x NVIDIA A100

80GB VRAM

Apple Silicon

2x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 2.6k · Context: 524K · Vocab: 250.6kx 48 layersRMSNormPre-AttentionGrouped-Query Attention32Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (8/100 experts)SwiGLUIntermediate: 768+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for K2 Horizon MoVA 36B A4B available.

Rankings

Overall Rank

-

Coding Rank

-

About K2 Horizon MoVA 36B A4B

K2 Horizon MoVA 36B A4B is a mixture-of-experts language model activating 4B parameters per token for efficient high-capacity inference. It is optimized for complex problem-solving, context retention, and reasoning tasks.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

32

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

10,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

2,560

Number of Layers

48

FFN Intermediate Size (Dense)

6,144

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

250,624

Mixture of Experts

Total Expert Parameters

4.0B

Number of Experts

100

Active Experts

8

Shared Experts

1

FFN Intermediate Size (per Expert)

768

Dense Layers Before MoE

-

About Kimi K2

Moonshot AI's Kimi K2 is a Mixture-of-Experts model featuring one trillion total parameters, activating 32 billion per token. Designed for agentic intelligence, it utilizes a sparse architecture with 384 experts and the MuonClip optimizer for training stability, supporting a 128K token context window.


Other Kimi K2 Models