Active Parameters
-
Context Length
1M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
-
Release Date
30 Jun 2026
Knowledge Cutoff
Jan 2026
No evaluation benchmarks for Claude Sonnet 5 available.
Overall Rank
-
Coding Rank
-
Anthropic's Claude Sonnet 5 is built for agentic workflows, long-horizon coding, and high-efficiency reasoning with a 1M token context window and adaptive thinking by default.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
-
Key-Value Heads
-
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
-
Activation Function
-
Dimensions
Hidden Dimension Size
-
Number of Layers
-
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Mixture of Experts
Total Expert Parameters
-
Number of Experts
-
Active Experts
-
Shared Experts
-
FFN Intermediate Size (per Expert)
-
Dense Layers Before MoE
-
Anthropic's Claude Sonnet 5, released June 30, 2026, is built to be the most agentic Sonnet-class model, offering near-Opus 4.8 level intelligence, coding, and tool use at Sonnet pricing. It features a 1M token context window by default, 128k maximum output tokens, and adaptive thinking enabled by default.
APX AI
Online