ApX logoApX logo

Phi-4

Parameters

14B

Context Length

16K

Modality

Text

Architecture

Dense

License

MIT License

Release Date

13 Dec 2024

Knowledge Cutoff

Nov 2024

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

31.08 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

16,000 tokens

33.65 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 3.1k · Context: 16K · Vocab: 100.4kx 40 layersRMSNormPre-AttentionGrouped-Query Attention24Q / 8KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwishIntermediate: 17.9k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#135

BenchmarkScoreRank

General Knowledge

MMLU

0.848

15

Professional Knowledge

MMLU Pro

0.7

63

General Text

Text Arena

1256

148

Rankings

Overall Rank

#135

Coding Rank

-

About Phi-4

Microsoft Phi-4 is a 14 billion parameter decoder-only Transformer model, developed as the latest iteration in Microsoft's series of small language models (SLMs). The model's primary objective is to deliver advanced reasoning capabilities efficiently, enabling deployment in environments with limited compute and memory, and for latency-sensitive applications. Phi-4 is designed to handle complex logical and mathematical tasks, along with general language processing, by focusing on the quality of its training data rather than solely on model scale.

A key innovation in Phi-4's architecture and training methodology lies in its strategic use of high-quality synthetic data, which constitutes a significant portion of its training corpus. This synthetic data, generated using techniques such as multi-agent prompting, instruction reversal, and self-revision workflows, is complemented by meticulously curated organic data from web content, academic books, and code repositories. This approach enables Phi-4 to acquire strong reasoning and problem-solving abilities, often surpassing models with larger parameter counts. The model's architecture retains a similar structure to its predecessor, Phi-3, but includes enhancements such as an extended context length.

Phi-4 supports a 16,000-token context length, allowing it to process and generate extensive long-form content. Its design prioritizes efficiency and robust performance in tasks requiring logical deduction, code generation, and scientific understanding. The model is intended for research and development, serving as a foundational component for generative AI features in various applications, particularly those demanding strong reasoning in resource-constrained or low-latency scenarios.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

24

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

250,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Hidden Dimension Size

3,072

Number of Layers

40

FFN Intermediate Size (Dense)

17,920

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

100,352

About Phi-4

The Microsoft Phi-4 model family comprises small language models prioritizing efficient, high-capability reasoning. Its development emphasizes robust data quality and sophisticated synthetic data integration. This approach enables enhanced performance and on-device deployment capabilities.


Other Phi-4 Models
Phi-4: Specifications and GPU VRAM Requirements