ApX logoApX logo

Ministral-8B-2410

Parameters

8B

Context Length

128K

Modality

Text

Architecture

Dense

License

Mistral Research License

Release Date

10 Oct 2024

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $0.10 · Output: $0.10

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

18.46 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

128,000 tokens

38.12 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 12.3k · Context: 128K · Vocab: 131.1kx 36 layersRMSNormPre-AttentionGrouped-Query Attention32Q / 8KV heads · SW: 32.8kHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwishIntermediate: 12.3k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#184

BenchmarkScoreRank

General Text

Text Arena

1237

149

General Knowledge

Reference
MMLU

0.65

29

Rankings

Overall Rank

#184

Coding Rank

-

About Ministral-8B-2410

Ministral-8B-2410 is a high-performance edge language model by Mistral AI engineered for local analytics, autonomous robotics, and on-device translation. It features interleaved sliding-window attention and function calling across a 128K context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

32

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

100,000,000

Sliding Window Attention

Yes

Sliding Window Size

32,768

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Hidden Dimension Size

12,288

Number of Layers

36

FFN Intermediate Size (Dense)

12,288

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

131,072

About Ministral

The Ministral model family, developed by Mistral AI, includes 3B and 8B parameter versions for on-device and edge computing. Designed for compute efficiency and low latency, these models support up to 128K context length. The 8B version incorporates an interleaved sliding-window attention pattern for efficient inference.


Other Ministral Models
Ministral-8B-2410: Specifications and GPU VRAM Requirements