Parameters
123B
Context Length
128K
Modality
Text
Architecture
Dense
License
Mistral Research License
Release Date
24 Jul 2024
Knowledge Cutoff
Mar 2024
API Pricing (per 1M)
Input: $2.00 · Output: $6.00
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
13x RTX 4090
24GB VRAM
Datacenter
4x NVIDIA A100
80GB VRAM
Apple Silicon
3x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
15x RTX 4090
24GB VRAM
Datacenter
5x NVIDIA A100
80GB VRAM
Apple Silicon
3x Apple M3 Max
128GB VRAM
Rank
#156
| Benchmark | Score | Rank |
|---|---|---|
General Text | 1314 | 134 |
Intelligence Index | 0.08 | 206 |
QA Assistant Archived | 0.964 | 5 |
StackEval Archived | 0.876 | 10 |
General Knowledge Reference | 0.84 | 15 |
Summarization Archived | 0.729 | 19 |
Overall Rank
#156
Coding Rank
-
Mistral Large 2 (2407) is Mistral AI's flagship 123B dense transformer model optimized for single-node enterprise inference and advanced code generation. It supports over 80 programming languages and precise function calling across a 128K context window.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
48
Key-Value Heads
8
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
12,288
Number of Layers
64
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Mistral Large 2 is a 123 billion parameter, dense transformer model engineered for advanced language and code generation, supporting over 80 programming languages. Its 128,000 token context window facilitates complex reasoning and long-context applications on a single node. Enhanced function calling capabilities are integrated.
Assistant
Online