Parameters
Undisclosed
Context Length
2M
Modality
Multimodal
Architecture
Undisclosed
License
Proprietary
Release Date
19 May 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $1.50 · Output: $9.00
Rank
#46
| Benchmark | Score | Rank |
|---|---|---|
StackUnseen | 0.846 | 11 |
Coding | high 0.78 | 22 |
General Text | high 1479 medium 1476 | 22 25 |
General | high 0.75 | 27 |
LiveBench Average | high 0.75 | 27 |
Mathematics | high 0.88 | 29 |
Agentic Coding | high 0.49 | 31 |
Reasoning | high 0.82 | 34 |
Agent Arena | high -3.48 medium -4.71 | 39 40 |
Web Development | high 1500 medium 1491 | 41 43 |
Data Analysis | high 0.65 | 43 |
Coding Index | high 0.70 | 46 |
Intelligence Index | high 0.33 medium 0.34 low 0.24 | 63 62 109 |
Agentic Index | high 0.27 | 63 |
Overall Rank
#46
Coding Rank
#35
Google's high-throughput, ultra-low-latency model debuted at Google I/O on May 19, 2026. Built directly for general availability scaling, it anchors Google's 24/7 proactive agentic infrastructure, dealing seamlessly with multimodal inputs and sprawling background tool operations.
Architecture specifications are undisclosed for proprietary models.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
Undisclosed
Key-Value Heads
Undisclosed
Attention Head Dimension
Undisclosed
Position Embedding
Absolute Position Embedding
RoPE Theta
Undisclosed
Sliding Window Attention
Undisclosed
Sliding Window Size
Undisclosed
Sliding Window Ratio
Undisclosed
Linear Attention
Undisclosed
Linear Attention Ratio
Undisclosed
Normalization
Undisclosed
Activation Function
Undisclosed
Dimensions
Hidden Dimension Size
Undisclosed
Number of Layers
Undisclosed
FFN Intermediate Size (Dense)
Undisclosed
Multi-Token Prediction Heads
Undisclosed
Tokenizer
Vocabulary Size
Undisclosed
The Gemini 3.5 generation represents Google's transition into highly proactive, autonomous agent ecosystems. Skipping standard preview staging to target instant production scalability, it features highly optimized structural modalities tailored for multi-tool execution pipelines, persistent background agent actions, and sub-second core text response latency.
APX AI
Online