ApX logoApX logo

Typhoon-2-8B

Parameters

8B

Context Length

128K

Modality

Text

Architecture

Dense

License

Apache-2.0

Release Date

1 Jun 2024

Knowledge Cutoff

Mar 2023

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

18.44 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

128,000 tokens

35.92 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 4.1k · Context: 128Kx 32 layersRMSNormPre-AttentionMulti-Head Attention32Q / 8KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwiGLU+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Typhoon-2-8B available.

Rankings

Overall Rank

-

Coding Rank

-

About Typhoon-2-8B

Typhoon-2-8B is a large language model specifically engineered to address the linguistic requirements of the Thai language while maintaining the broad capabilities of the Llama 3 architecture. Developed by SCB 10X, the model undergoes a specialized training process that involves extending the base tokenizer with Thai-specific tokens and performing continual pre-training on a high-quality Thai corpus. This adaptation ensures that the model can process Thai text with higher efficiency and accuracy compared to general-purpose multilingual models, particularly in domains such as Thai law, local administration, and cultural contexts.

The technical architecture follows a dense transformer structure utilizing Grouped-Query Attention (GQA) to optimize inference speed and memory consumption. It incorporates Rotary Positional Embeddings (RoPE) and is configured with a context window of 128,000 tokens, enabling the processing of long-form documents and complex multi-turn conversations. The model utilizes the SwiGLU activation function and Root Mean Square Layer Normalization (RMSNorm) to stabilize training and improve representation learning across its 32 layers.

Function calling capabilities are integrated into the model, allowing it to interact with external tools and APIs by generating structured data outputs. This functionality makes it suitable for agentic workflows, automated administrative tasks, and specialized information retrieval systems where precise Thai language understanding is required. The model is released under the Apache 2.0 license, facilitating both research and commercial applications in the Thai technology ecosystem.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

32

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

Absolute Position Embedding

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

4,096

Number of Layers

32

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

About Typhoon

Typhoon is a Thai language model family developed by SCB 10X. It is specifically optimized for the Thai language, addressing complexities such as the lack of word delimiters and tonal nuances. The models are trained on Thai-centric datasets including legal, cultural, and historical documents to ensure localized context and knowledge.


Other Typhoon Models