ApX logoApX logo

Mistral-Small-2501

Parameters

24B

Context Length

33K

Modality

Text

Architecture

Dense

License

Apache 2.0

Release Date

13 Jan 2025

Knowledge Cutoff

Oct 2023

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

52.03 GB VRAM

Consumer

3x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

32,768 tokens

56.13 GB VRAM

Consumer

3x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 32.8k · Context: 33K · Vocab: 131.1kx 40 layersRMSNormPre-AttentionGrouped-Query Attention24Q / 6KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwiGLUIntermediate: 32.8k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#127

BenchmarkScoreRank

0.913

13

0.747

18

General Knowledge

MMLU

0.81

20

0.12

32

General Text

Text Arena

1274

145

Intelligence Index

Artificial Analysis

0.07

244

Rankings

Overall Rank

#127

Coding Rank

#139

About Mistral-Small-2501

Mistral Small 3 (2501) is a 24B-parameter dense model from Mistral AI engineered for rapid response times and high-throughput enterprise tasks. Released under Apache 2.0, it offers low-latency reasoning and function calling across a 32K context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

24

Key-Value Heads

6

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

100,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

32,768

Number of Layers

40

FFN Intermediate Size (Dense)

32,768

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

131,072

About Mistral Small 3

Mistral Small 3, a 24 billion parameter model, was designed for efficient, low-latency generative AI tasks. Its optimized architecture supports local deployment and includes multimodal understanding, multilingual capabilities, and a 128,000-token context window.


Other Mistral Small 3 Models
  • No related models available
Mistral-Small-2501: Specifications and GPU VRAM Requirements