ApX logoApX logo

Sarvam-105B

Active Parameters

106B

Context Length

128K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

6 Mar 2026

Knowledge Cutoff

-

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

224.73 GB VRAM

Consumer

11x RTX 4090

24GB VRAM

Datacenter

4x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

128,000 tokens

303.37 GB VRAM

Consumer

16x RTX 4090

24GB VRAM

Datacenter

5x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 4.1k · Context: 128K · Vocab: 262.1kx 32 layersRMSNormPre-AttentionMulti-Layer Attention64 headsHead dim: 576+RMSNormPre-FFNSparse MoE FFN (8/128 experts)SwiGLUIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Sarvam-105B available.

Rankings

Overall Rank

-

Coding Rank

-

About Sarvam-105B

Sarvam-105B is an advanced Mixture-of-Experts (MoE) model with 106B total parameters and 10.3B active parameters, designed for superior performance across complex tasks. Released March 6, 2026 under Apache 2.0 license. Uses MLA-style attention stack with decoupled QK head dimensions (q_head_dim=192, v_head_dim=128), large head_dim of 576, and 128 experts with top-8 routing. Features 128K native context (extensible via YaRN scaling with factor 40), and delivers exceptional performance in agentic tasks, mathematics, and coding. Consistently matches or surpasses major closed-source models with state-of-the-art results across 22 Indian languages while maintaining competitive global benchmark performance.

Technical Specifications

Attention

Attention Structure

Multi-Layer Attention

Attention Heads

64

Key-Value Heads

-

Attention Head Dimension

576

Position Embedding

ROPE

RoPE Theta

10,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

4,096

Number of Layers

32

FFN Intermediate Size (Dense)

2,048

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

262,144

Mixture of Experts

Total Expert Parameters

10.3B

Number of Experts

128

Active Experts

8

Shared Experts

1

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

1

About Sarvam

Sarvam AI's sovereign foundation models built for India's languages, culture, and context. Released in March 2026, these advanced Mixture-of-Experts (MoE) models offer state-of-the-art performance across 22 Indian languages while maintaining competitive results on global benchmarks. Designed with focus on reasoning, coding, multilingual capabilities, and agentic tasks. Open-sourced under Apache 2.0 license, optimized for practical deployment from resource-constrained environments to high-performance applications.


Other Sarvam Models