Parameters
8B
Context Length
128K
Modality
Text
Architecture
Dense
License
Mistral Research License
Release Date
10 Oct 2024
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $0.10 · Output: $0.10
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
1x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
2x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
Rank
#184
| Benchmark | Score | Rank |
|---|---|---|
General Text | 1237 | 149 |
General Knowledge Reference | 0.65 | 29 |
Overall Rank
#184
Coding Rank
-
Ministral-8B-2410 is a high-performance edge language model by Mistral AI engineered for local analytics, autonomous robotics, and on-device translation. It features interleaved sliding-window attention and function calling across a 128K context window.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
32
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
ROPE
RoPE Theta
100,000,000
Sliding Window Attention
Yes
Sliding Window Size
32,768
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
Swish
Dimensions
Hidden Dimension Size
12,288
Number of Layers
36
FFN Intermediate Size (Dense)
12,288
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
131,072
The Ministral model family, developed by Mistral AI, includes 3B and 8B parameter versions for on-device and edge computing. Designed for compute efficiency and low latency, these models support up to 128K context length. The 8B version incorporates an interleaved sliding-window attention pattern for efficient inference.
APX AI
Online