ApX logoApX logo

Qwen3 Next 80B A3B

Active Parameters

80B

Context Length

66K

Modality

Reasoning

Architecture

Mixture of Experts (MoE)

License

Apache-2.0

Release Date

1 Feb 2026

Knowledge Cutoff

Jun 2025

API Pricing (per 1M)

Input: $0.15 · Output: $1.20

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

169.61 GB VRAM

Consumer

9x RTX 4090

24GB VRAM

Datacenter

3x NVIDIA A100

80GB VRAM

Apple Silicon

2x Apple M3 Max

128GB VRAM

66,000 tokens

176.31 GB VRAM

Consumer

9x RTX 4090

24GB VRAM

Datacenter

3x NVIDIA A100

80GB VRAM

Apple Silicon

2x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 2k · Context: 66K · Vocab: 151.9kx 48 layersRMSNormPre-AttentionMulti-Head Attention16Q / 2KV headsHead dim: 256+RMSNormPre-FFNSparse MoE FFN (10/512 experts)SwiGLUIntermediate: 512+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#110

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.806

21

Graduate-Level QA

GPQA

0.772

58

General Text

Text Arena

1399

87

Agentic Index

Artificial Analysis

0.02

101

0.17

118

Intelligence Index

Artificial Analysis
auto

0.11

Standard

0.10

187

196

Rankings

Overall Rank

#110

Coding Rank

#113

About Qwen3 Next 80B A3B

Qwen3-Next-80B-A3B is a sparse Mixture-of-Experts model by Alibaba activating 3B parameters via hybrid Gated DeltaNet linear attention. It delivers ultra-fast sequence modeling and complex mathematical reasoning over an extensible 1M context window.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

16

Key-Value Heads

2

Attention Head Dimension

256

Position Embedding

Absolute Position Embedding

RoPE Theta

10,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

2,048

Number of Layers

48

FFN Intermediate Size (Dense)

512

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

151,936

Mixture of Experts

Total Expert Parameters

79.0B

Number of Experts

512

Active Experts

10

Shared Experts

-

FFN Intermediate Size (per Expert)

512

Dense Layers Before MoE

-

About Qwen 3

The Alibaba Qwen 3 model family comprises dense and Mixture-of-Experts (MoE) architectures, with parameter counts from 0.6B to 235B. Key innovations include a hybrid reasoning system, offering 'thinking' and 'non-thinking' modes for adaptive processing, and support for extensive context windows, enhancing efficiency and scalability.


Other Qwen 3 Models
Qwen3 Next 80B A3B: Specifications and GPU VRAM Requirements