ApX logoApX logo

Sahabat-AI-Gemma2-9B-Instruct

Parameters

9.2B

Context Length

8K

Modality

Text

Architecture

Dense

License

Gemma-Community

Release Date

14 Nov 2024

Knowledge Cutoff

Mar 2024

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

21.19 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

8,192 tokens

23.78 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 3.6k · Context: 8K · Vocab: 256kx 42 layersRMSNormPre-AttentionMulti-Head Attention16Q / 8KV heads · SW: 4.1kHead dim: 256+RMSNormPre-FFNFeed-Forward NetworkGated GELUIntermediate: 14.3k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Sahabat-AI-Gemma2-9B-Instruct available.

Rankings

Overall Rank

-

Coding Rank

-

About Sahabat-AI-Gemma2-9B-Instruct

Sahabat-AI-Gemma2-9B-Instruct is a specialized large language model developed through a strategic collaboration between GoTo Group, Indosat Ooredoo Hutchison, and AI Singapore. Built upon the Google Gemma 2 architecture, this variant is the result of continued pre-training (CPT) and intensive instruction tuning specifically tailored for the Indonesian linguistic ecosystem. It is engineered to provide high-fidelity conversational capabilities not only in standard Bahasa Indonesia but also in major regional dialects, including Javanese and Sundanese, addressing the cultural and linguistic nuances inherent to the Indonesian archipelago.

The underlying architecture follows a decoder-only transformer design that incorporates several modern refinements for efficiency and stability. It utilizes Grouped-Query Attention (GQA) to optimize inference throughput and memory bandwidth, which is particularly effective for maintaining performance during long-context processing. For training stability and representational accuracy, the model employs RMSNorm for pre- and post-normalization across layers and integrates logit soft-capping to prevent divergence. The instruction-tuning phase involved a supervised fine-tuning process using a localized dataset of over 600,000 instruction-completion pairs, followed by on-policy alignment and model merging to refine its response quality and adherence to complex prompts.

Technically, the model is optimized for a wide array of natural language processing tasks, including sentiment analysis, toxicity detection, causal reasoning, and abstractive summarization within Southeast Asian contexts. By leveraging the base Gemma 2 9B weights, it inherits a robust world-knowledge foundation while specializing in regional idioms and cultural contexts that are often underrepresented in global models. This makes it a suitable candidate for developers building localized digital assistants, automated customer service interfaces, and educational tools designed for the Indonesian market.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

16

Key-Value Heads

8

Attention Head Dimension

256

Position Embedding

Absolute Position Embedding

RoPE Theta

10,000

Sliding Window Attention

Yes

Sliding Window Size

4,096

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Gated GELU

Dimensions

Hidden Dimension Size

3,584

Number of Layers

42

FFN Intermediate Size (Dense)

14,336

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

256,000

About Sahabat-AI

Sahabat-AI is an Indonesian language model family co-initiated by GoTo and Indosat Ooredoo Hutchison. Developed with AI Singapore and NVIDIA, it is a collection of models (based on Gemma 2 and Llama 3) specifically optimized for Bahasa Indonesia and regional languages like Javanese and Sundanese.


Other Sahabat-AI Models