Parameters
32B
Context Length
131K
Modality
Text
Architecture
Dense
License
Apache 2.0
Release Date
19 Sept 2024
Knowledge Cutoff
Mar 2024
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
4x RTX 4090
24GB VRAM
Datacenter
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
131,072 tokens
Consumer
5x RTX 4090
24GB VRAM
Datacenter
2x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
Rank
#63
| Benchmark | Score | Rank |
|---|---|---|
General Knowledge | 0.833 | 18 |
Overall Rank
#63
Coding Rank
-
The Qwen2.5-32B model is a significant component of the Qwen2.5 series of large language models, developed by the Qwen team at Alibaba Cloud. This iteration builds upon its predecessors by offering enhanced capabilities for a broad spectrum of natural language processing tasks. Its design prioritizes robust instruction following, effective long-text generation, and sophisticated comprehension and production of structured data, including JSON formats. The model also demonstrates improved stability when confronted with diverse system prompts, which is advantageous for developing conversational agents and setting specific dialogue conditions. Furthermore, it provides comprehensive multilingual support across more than 29 languages, expanding its applicability in global contexts.
Architecturally, Qwen2.5-32B is a dense, decoder-only transformer model. It integrates several advanced components to optimize performance and efficiency. These include Rotary Position Embeddings (RoPE) for effective positional encoding, SwiGLU as the activation function for enhanced non-linearity, and RMSNorm for stable training and improved convergence. To optimize inference speed and Key-Value cache utilization, the model employs Grouped Query Attention (GQA). The underlying training regimen involved a massive dataset, expanded to approximately 18 trillion tokens, which contributed to its enriched knowledge base, particularly in domains such as coding, mathematics, and various languages.
The operational characteristics of Qwen2.5-32B demonstrate notable performance across various complex tasks. This model variant is adept at handling extended contexts, supporting sequences up to 131,072 tokens. Its ability to generate long texts, with outputs extending up to 8,192 tokens, makes it suitable for applications requiring detailed responses or extensive content creation. While the base model is general-purpose, the architectural foundations of Qwen2.5 have also been utilized in specialized variants, such as those optimized for coding or multimodal vision-language tasks, underscoring the versatility of the Qwen2.5 framework.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
96
Key-Value Heads
8
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
1,000,000
Sliding Window Attention
No
Sliding Window Size
131,072
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
8,192
Number of Layers
60
FFN Intermediate Size (Dense)
27,648
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
152,064
Qwen2.5 by Alibaba is a family of dense, decoder-only language models available in various sizes, with some variants utilizing Mixture-of-Experts. These models are pretrained on large-scale datasets, supporting extended context lengths and multilingual communication. The family includes specialized models for coding, mathematics, and multimodal tasks, such as vision and audio processing.
APX AI
Online