ApX logoApX logo

Qwen 3.8 27B

Parameters

27B

Context Length

262K

Modality

Text

Architecture

Dense

License

Apache 2.0

Release Date

5 Aug 2026

Knowledge Cutoff

-

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

58.48 GB VRAM

Consumer

3x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

262,144 tokens

130.36 GB VRAM

Consumer

7x RTX 4090

24GB VRAM

Datacenter

2x NVIDIA A100

80GB VRAM

Apple Silicon

2x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 5.1k · Context: 262K · Vocab: 248.3kx 64 layersRMSNormPre-AttentionGrouped-Query Attention24Q / 4KV headsHead dim: 256+RMSNormPre-FFNFeed-Forward NetworkSwiGLUIntermediate: 17.4k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#55

BenchmarkScoreRank

Agentic Coding

LiveBench Agentic

0.61

9

Web Development

WebDev Arena

1598

15

0.77

26

0.75

27

LiveBench Average

LiveBench Average

0.75

33

0.76

42

0.80

42

Agent Arena

Agent Arena

0.01

45

0.86

52

General Text

Text Arena

1435

89

Rankings

Overall Rank

#55

Coding Rank

#43

About Qwen 3.8 27B

Qwen 3.8 27B is a 27-billion parameter dense foundation model by Alibaba that offers balanced efficiency and strong reasoning performance. It excels at multi-turn conversational tasks, programming, and long-context understanding.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

24

Key-Value Heads

4

Attention Head Dimension

256

Position Embedding

ROPE

RoPE Theta

10,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

Yes

Linear Attention Ratio

75.0%

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

5,120

Number of Layers

64

FFN Intermediate Size (Dense)

17,408

Multi-Token Prediction Heads

1

Tokenizer

Vocabulary Size

248,320

About Qwen 3.8

Alibaba's Qwen 3.8 generation represents the frontier hybrid Mixture-of-Experts architecture designed for coding, professional work, research, and long-horizon agentic tasks. It features a 2.4-trillion parameter architecture (95B active per token) combining Gated DeltaNet linear attention with standard Gated Attention, available both as open weights and as a hosted flagship service.


Other Qwen 3.8 Models
Qwen 3.8 27B: Specifications and GPU VRAM Requirements