ApX logoApX logo

Sarvam-30B

Active Parameters

32B

Context Length

128K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

6 Mar 2026

Knowledge Cutoff

-

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

68.72 GB VRAM

Consumer

4x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

128,000 tokens

71.31 GB VRAM

Consumer

4x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 4.1k · Context: 128K · Vocab: 262.1kx 19 layersRMSNormPre-AttentionGrouped-Query Attention64Q / 4KV headsHead dim: 64+RMSNormPre-FFNSparse MoE FFN (6/128 experts)SwiGLUIntermediate: 1k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Sarvam-30B available.

Rankings

Overall Rank

-

Coding Rank

-

About Sarvam-30B

Sarvam-30B is an advanced Mixture-of-Experts (MoE) model with 32B total parameters and 2.4B active parameters, designed for practical deployment in resource-constrained environments. Released March 6, 2026 under Apache 2.0 license. Uses 19 layers with 128 experts, top-6 routing, grouped KV attention (4 heads), and extremely high rope_theta (8e6) for long-context stability. Delivers state-of-the-art performance across 22 Indian languages with strong reasoning, reliable coding ability, and best-in-class conversational quality. Optimized for multilingual voice calls with tool calling capabilities, throughput, and memory efficiency.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

64

Key-Value Heads

4

Attention Head Dimension

64

Position Embedding

ROPE

RoPE Theta

8,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

4,096

Number of Layers

19

FFN Intermediate Size (Dense)

1,024

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

262,144

Mixture of Experts

Total Expert Parameters

2.4B

Number of Experts

128

Active Experts

6

Shared Experts

1

FFN Intermediate Size (per Expert)

1,024

Dense Layers Before MoE

1

About Sarvam

Sarvam AI's sovereign foundation models built for India's languages, culture, and context. Released in March 2026, these advanced Mixture-of-Experts (MoE) models offer state-of-the-art performance across 22 Indian languages while maintaining competitive results on global benchmarks. Designed with focus on reasoning, coding, multilingual capabilities, and agentic tasks. Open-sourced under Apache 2.0 license, optimized for practical deployment from resource-constrained environments to high-performance applications.


Other Sarvam Models