ApX logoApX logo

Mixtral-8x7B-v0.1

Active Parameters

46.7B

Context Length

33K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

9 Dec 2023

Knowledge Cutoff

Nov 2022

API Pricing (per 1M)

Input: $0.45 · Output: $0.70

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

99.71 GB VRAM

Consumer

5x RTX 4090

24GB VRAM

Datacenter

2x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

32,768 tokens

104.08 GB VRAM

Consumer

5x RTX 4090

24GB VRAM

Datacenter

2x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 4.1k · Context: 33K · Vocab: 32kx 32 layersRMSNormPre-AttentionGrouped-Query Attention32Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (2/8 experts)SwishIntermediate: 14.3k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#181

BenchmarkScoreRank

General Text

Text Arena

1197

160

Intelligence Index

Artificial Analysis

0.05

228

Rankings

Overall Rank

#181

Coding Rank

-

About Mixtral-8x7B-v0.1

Mixtral-8x7B-v0.1 is an open Sparse Mixture-of-Experts model from Mistral AI activating 12.9B parameters per token for high-efficiency inference. It delivers strong multilingual and coding performance across a 32K token context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

32

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

1,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Hidden Dimension Size

4,096

Number of Layers

32

FFN Intermediate Size (Dense)

14,336

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

32,000

Mixture of Experts

Total Expert Parameters

7.0B

Number of Experts

8

Active Experts

2

Shared Experts

-

FFN Intermediate Size (per Expert)

14,336

Dense Layers Before MoE

-

About Mixtral

The Mixtral model family, developed by Mistral AI, employs a sparse Mixture-of-Experts (SMoE) architecture. This design utilizes multiple expert networks per layer, where a router selects a subset to process each token. This enables large total parameter counts while maintaining computational efficiency by activating only a fraction of parameters per forward pass.


Other Mixtral Models