Parameters
Undisclosed
Context Length
1.05M
Modality
Multimodal
Architecture
Undisclosed
License
Proprietary
Release Date
21 Jul 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $0.75 · Output: $3.75
Rank
#94
| Benchmark | Score | Rank |
|---|---|---|
General Text | high 1480 | 22 |
Coding | high 0.78 | 26 |
Reasoning | high 0.85 | 29 |
Web Development | high 1537 | 31 |
LiveBench Average | high 0.74 | 33 |
General | high 0.74 | 34 |
Mathematics | high 0.86 | 37 |
Agentic Coding | high 0.43 | 41 |
Data Analysis | high 0.63 | 44 |
Agent Arena | high -5.94 | 44 |
Coding Index | high 0.69 | 48 |
Intelligence Index | high 0.34 | 54 |
Agentic Index | high 0.30 | 60 |
Overall Rank
#94
Coding Rank
#50
Gemini 3.6 Flash is a high-efficiency multimodal model from Google optimized for rapid coding, web development, and agentic execution. It delivers high-precision outputs with minimal latency over a 1M token context window.
Architecture specifications are undisclosed for proprietary models.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
Undisclosed
Key-Value Heads
Undisclosed
Attention Head Dimension
Undisclosed
Position Embedding
Absolute Position Embedding
RoPE Theta
Undisclosed
Sliding Window Attention
Undisclosed
Sliding Window Size
Undisclosed
Sliding Window Ratio
Undisclosed
Linear Attention
Undisclosed
Linear Attention Ratio
Undisclosed
Normalization
Undisclosed
Activation Function
Undisclosed
Dimensions
Hidden Dimension Size
Undisclosed
Number of Layers
Undisclosed
FFN Intermediate Size (Dense)
Undisclosed
Multi-Token Prediction Heads
Undisclosed
Tokenizer
Vocabulary Size
Undisclosed
Google's Gemini 3.6 generation is optimized for high-efficiency multimodal processing, rapid agentic tool execution, and low-latency coding workflows over an extended context window.
Assistant
Online