ApX logoApX logo

Phi-4

Parameters

14B

Context Length

16K

Modality

Text

Architecture

Dense

License

MIT License

Release Date

13 Dec 2024

Knowledge Cutoff

Nov 2024

API Pricing (per 1M)

Input: $0.13 · Output: $0.50

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

31.08 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

16,000 tokens

33.65 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 3.1k · Context: 16K · Vocab: 100.4kx 40 layersRMSNormPre-AttentionGrouped-Query Attention24Q / 8KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwishIntermediate: 17.9k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#149

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.704

40

Graduate-Level QA

GPQA

0.561

99

General Text

Text Arena

1256

154

Intelligence Index

Artificial Analysis

0.06

276

General Knowledge

Reference
MMLU

0.848

14

Rankings

Overall Rank

#149

Coding Rank

-

About Phi-4

Phi-4 is Microsoft's 14B parameter dense transformer model engineered with synthetic reasoning data for complex mathematics, logic, and scientific problem-solving. It delivers high-accuracy reasoning across a 16K token context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

24

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

250,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

3,072

Number of Layers

40

FFN Intermediate Size (Dense)

17,920

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

100,352

About Phi-4

The Microsoft Phi-4 model family comprises small language models prioritizing efficient, high-capability reasoning. Its development emphasizes robust data quality and sophisticated synthetic data integration. This approach enables enhanced performance and on-device deployment capabilities.


Other Phi-4 Models
Phi-4: Specifications and GPU VRAM Requirements