ApX logoApX logo

Mistral-Large-2407

Parameters

123B

Context Length

128K

Modality

Text

Architecture

Dense

License

Mistral Research License

Release Date

24 Jul 2024

Knowledge Cutoff

Mar 2024

API Pricing (per 1M)

Input: $2.00 · Output: $6.00

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

260.08 GB VRAM

Consumer

13x RTX 4090

24GB VRAM

Datacenter

4x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

128,000 tokens

295.03 GB VRAM

Consumer

15x RTX 4090

24GB VRAM

Datacenter

5x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 12.3k · Context: 128Kx 64 layersRMSNormPre-AttentionGrouped-Query Attention48Q / 8KV headsHead dim: 256+RMSNormPre-FFNFeed-Forward NetworkSwiGLU+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#156

BenchmarkScoreRank

General Text

Text Arena

1314

134

Intelligence Index

Artificial Analysis

0.08

206

QA Assistant

Archived
ProLLM QA Assistant

0.964

5

StackEval

Archived
ProLLM Stack Eval

0.876

10

General Knowledge

Reference
MMLU

0.84

15

Summarization

Archived
ProLLM Summarization

0.729

19

Rankings

Overall Rank

#156

Coding Rank

-

About Mistral-Large-2407

Mistral Large 2 (2407) is Mistral AI's flagship 123B dense transformer model optimized for single-node enterprise inference and advanced code generation. It supports over 80 programming languages and precise function calling across a 128K context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

48

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

12,288

Number of Layers

64

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

About Mistral Large 2

Mistral Large 2 is a 123 billion parameter, dense transformer model engineered for advanced language and code generation, supporting over 80 programming languages. Its 128,000 token context window facilitates complex reasoning and long-context applications on a single node. Enhanced function calling capabilities are integrated.


Other Mistral Large 2 Models
  • No related models available