Active Parameters
389B
Context Length
28K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
Tencent Hunyuan Community License
Release Date
5 Nov 2024
Knowledge Cutoff
Sep 2024
API Pricing (per 1M)
Self-hosted only
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
46x RTX 4090
24GB VRAM
Datacenter
12x NVIDIA A100
80GB VRAM
Apple Silicon
10x Apple M3 Max
128GB VRAM
28,000 tokens
Consumer
50x RTX 4090
24GB VRAM
Datacenter
13x NVIDIA A100
80GB VRAM
Apple Silicon
10x Apple M3 Max
128GB VRAM
Rank
#127
| Benchmark | Score | Rank |
|---|---|---|
General Text | 1326 | 133 |
Overall Rank
#127
Coding Rank
-
Hunyuan-DiT is Tencent's large-scale Mixture-of-Experts diffusion transformer engineered for high-fidelity text-to-image synthesis up to 4096x4096 resolution. It features bilingual CLIP and T5 encoders for fine-grained prompt comprehension and multi-turn editing.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
64
Key-Value Heads
64
Attention Head Dimension
-
Position Embedding
Absolute Position Embedding
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
Layer Normalization
Activation Function
GELU
Dimensions
Hidden Dimension Size
4,096
Number of Layers
60
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Mixture of Experts
Total Expert Parameters
52.0B
Number of Experts
32
Active Experts
2
Shared Experts
-
FFN Intermediate Size (per Expert)
-
Dense Layers Before MoE
-
Tencent Hunyuan large language models with various capabilities.
APX AI
Online