ApX logoApX logo

K2 Horizon 375B A23B

Active Parameters

375B

Context Length

524K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

1 Sept 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Self-hosted only

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

789.27 GB VRAM

Consumer

44x RTX 4090

24GB VRAM

Datacenter

12x NVIDIA A100

80GB VRAM

Apple Silicon

9x Apple M3 Max

128GB VRAM

524,288 tokens

926.55 GB VRAM

Consumer

53x RTX 4090

24GB VRAM

Datacenter

14x NVIDIA A100

80GB VRAM

Apple Silicon

11x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 6.1k · Context: 524K · Vocab: 250.6kx 61 layersRMSNormPre-AttentionGrouped-Query Attention48Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (8/192 experts)SwiGLUIntermediate: 1.8k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for K2 Horizon 375B A23B available.

Rankings

Overall Rank

-

Coding Rank

-

About K2 Horizon 375B A23B

K2 Horizon 375B A23B is a massive Mixture-of-Experts language model activating 23B parameters per token. It is built for demanding reasoning, advanced code synthesis, and deep multi-turn conversational agents.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

48

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

10,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

6,144

Number of Layers

61

FFN Intermediate Size (Dense)

16,384

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

250,624

Mixture of Experts

Total Expert Parameters

23.0B

Number of Experts

192

Active Experts

8

Shared Experts

1

FFN Intermediate Size (per Expert)

1,792

Dense Layers Before MoE

-

About Kimi K2

Moonshot AI's Kimi K2 is a Mixture-of-Experts model featuring one trillion total parameters, activating 32 billion per token. Designed for agentic intelligence, it utilizes a sparse architecture with 384 experts and the MuonClip optimizer for training stability, supporting a 128K token context window.


Other Kimi K2 Models