Active Parameters
80B
Context Length
262K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
Apache 2.0
Release Date
12 Sept 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Self-hosted only
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
9x RTX 4090
24GB VRAM
Datacenter
3x NVIDIA A100
80GB VRAM
Apple Silicon
2x Apple M3 Max
128GB VRAM
262,144 tokens
Consumer
10x RTX 4090
24GB VRAM
Datacenter
3x NVIDIA A100
80GB VRAM
Apple Silicon
2x Apple M3 Max
128GB VRAM
No evaluation benchmarks for AliceAI Foundation 80B A3B Base available.
Overall Rank
-
Coding Rank
-
AliceAI Foundation 80B A3B Base is a sparse mixture-of-experts foundation model developed by Yandex, routing tokens through 3B active parameters out of an 80B total parameter pool. It provides strong multilingual and reasoning abilities with high computational efficiency during inference.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
16
Key-Value Heads
2
Attention Head Dimension
256
Position Embedding
ROPE
RoPE Theta
1,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
75.0%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Auxiliary Parameters
-
Hidden Dimension Size
2,048
Number of Layers
48
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
129,024
Mixture of Experts
Total Expert Parameters
3.0B
Number of Experts
512
Active Experts
10
Shared Experts
1
FFN Intermediate Size (per Expert)
512
Dense Layers Before MoE
0
The AliceAI model family developed by Yandex.
Assistant
Online