Parameters
4B
Context Length
1.05M
Modality
Text
Architecture
Dense
License
Apache 2.0
Release Date
24 Aug 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Self-hosted only
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
1x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
9x RTX 4090
24GB VRAM
Datacenter
3x NVIDIA A100
80GB VRAM
Apple Silicon
2x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Spark X2.5 4B available.
Overall Rank
-
Coding Rank
-
Spark X2.5 4B is a compact, high-efficiency language model designed for edge deployment and fast local inference. It delivers strong multilingual and reasoning performance relative to its lightweight parameter scale.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
16
Key-Value Heads
4
Attention Head Dimension
256
Position Embedding
ROPE
RoPE Theta
5,000,000
Sliding Window Attention
Yes
Sliding Window Size
512
Sliding Window Ratio
75.0%
Linear Attention
No
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
GELU
Dimensions
Hidden Dimension Size
2,560
Number of Layers
36
FFN Intermediate Size (Dense)
10,240
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
131,072
The Muse Spark model family developed by Thinking Machine Labs.
APX AI
Online