ApX logoApX logo

GPT-OSS 20B

Active Parameters

21B

Context Length

128K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

5 Aug 2025

Knowledge Cutoff

Jun 2024

API Pricing (per 1M)

Input: $0.06 · Output: $0.19

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

45.65 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

128,000 tokens

52.21 GB VRAM

Consumer

3x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 2.9k · Context: 128K · Vocab: 201.1kx 24 layersRMSNormPre-AttentionMulti-Head Attention64Q / 8KV heads · SW: 128Head dim: 64+RMSNormPre-FFNSparse MoE FFN (4/32 experts)SwiGLU+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#115

BenchmarkScoreRank

Graduate-Level QA

GPQA
high

0.742

Standard

0.715

77

79

Agentic Index

Artificial Analysis
high

0.01

111

General Text

Text Arena

1317

128

high

0.21

132

Intelligence Index

Artificial Analysis
high

0.09

low

0.10

187

178

Summarization

Archived
ProLLM Summarization

0.863

6

General Knowledge

Reference
MMLU

0.853

11

Rankings

Overall Rank

#115

Coding Rank

#118

About GPT-OSS 20B

GPT-OSS 20B is an open-weight Mixture-of-Experts language model by OpenAI designed for efficient local reasoning on consumer hardware. It activates 3.6B parameters per token and features native tool use across a 128K context window.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

64

Key-Value Heads

8

Attention Head Dimension

64

Position Embedding

Absolute Position Embedding

RoPE Theta

150,000

Sliding Window Attention

Yes

Sliding Window Size

128

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

2,880

Number of Layers

24

FFN Intermediate Size (Dense)

2,880

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

201,088

Mixture of Experts

Total Expert Parameters

3.6B

Number of Experts

32

Active Experts

4

Shared Experts

-

FFN Intermediate Size (per Expert)

-

Dense Layers Before MoE

-

About GPT-OSS

Open-weight language models from OpenAI.


Other GPT-OSS Models