ApX logoApX logo

Mixtral-8x22B-v0.1

Active Parameters

176B

Context Length

66K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

10 Apr 2024

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $0.90 · Output: $0.90

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

371.35 GB VRAM

Consumer

19x RTX 4090

24GB VRAM

Datacenter

6x NVIDIA A100

80GB VRAM

Apple Silicon

4x Apple M3 Max

128GB VRAM

65,536 tokens

386.88 GB VRAM

Consumer

20x RTX 4090

24GB VRAM

Datacenter

6x NVIDIA A100

80GB VRAM

Apple Silicon

4x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 1k · Context: 66K · Vocab: 32kx 56 layersRMSNormPre-AttentionGrouped-Query Attention48Q / 8KV headsHead dim: 21+RMSNormPre-FFNSparse MoE FFN (2/8 experts)SwiGLUIntermediate: 16.4k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#191

BenchmarkScoreRank

General Text

Text Arena

1229

166

Intelligence Index

Artificial Analysis

0.06

240

Summarization

Archived
ProLLM Summarization

0.587

25

Rankings

Overall Rank

#191

Coding Rank

-

About Mixtral-8x22B-v0.1

Mixtral-8x22B-v0.1 is Mistral AI's 176B Sparse Mixture-of-Experts model activating 39B parameters for complex mathematical, coding, and multilingual tasks. It features native function calling and high computational throughput across a 64K context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

48

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

1,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

1,024

Number of Layers

56

FFN Intermediate Size (Dense)

16,384

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

32,000

Mixture of Experts

Total Expert Parameters

22.0B

Number of Experts

8

Active Experts

2

Shared Experts

-

FFN Intermediate Size (per Expert)

16,384

Dense Layers Before MoE

-

About Mixtral

The Mixtral model family, developed by Mistral AI, employs a sparse Mixture-of-Experts (SMoE) architecture. This design utilizes multiple expert networks per layer, where a router selects a subset to process each token. This enables large total parameter counts while maintaining computational efficiency by activating only a fraction of parameters per forward pass.


Other Mixtral Models
Mixtral-8x22B-v0.1: Specifications and GPU VRAM Requirements