Parameters
-
Context Length
32K
Modality
Text
Architecture
Dense
License
Apache 2.0
Release Date
15 Jan 2025
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
1x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
32,000 tokens
Consumer
1x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
Rank
#150
| Benchmark | Score | Rank |
|---|---|---|
Coding | 0.11 | 34 |
Overall Rank
#150
Coding Rank
#86
Codestral 25.01 is Mistral AI's specialized coding model with deep understanding of software development. Features enhanced capabilities for code generation, completion, debugging, and refactoring across multiple programming languages. Trained on diverse codebases with focus on modern development practices, design patterns, and code quality. Excels at understanding developer intent and generating idiomatic, well-structured code. January 2025 release brings improved accuracy and expanded language support.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
48
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
Absolute Position Embedding
RoPE Theta
1,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
6,144
Number of Layers
56
FFN Intermediate Size (Dense)
16,384
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
32,768
Codestral is a Mistral AI model designed for code generation and comprehension. It supports over 80 programming languages. The model family includes a 22 billion parameter variant.
APX AI
Online