ApX logoApX logo

Gemini 3.5 Flash

Parameters

Undisclosed

Context Length

2M

Modality

Multimodal

Architecture

Undisclosed

License

Proprietary

Release Date

19 May 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $1.50 · Output: $9.00

Evaluation Benchmarks

Rank

#46

BenchmarkScoreRank

0.846

11

high

0.78

22

General Text

Text Arena
high

1479

medium

1476

22

25

high

0.75

27

LiveBench Average

LiveBench Average
high

0.75

27

high

0.88

29

Agentic Coding

LiveBench Agentic
high

0.49

31

high

0.82

34

Agent Arena

Agent Arena
high

-3.48

medium

-4.71

39

40

Web Development

WebDev Arena
high

1500

medium

1491

41

43

high

0.65

43

high

0.70

46

Intelligence Index

Artificial Analysis
high

0.33

medium

0.34

low

0.24

63

62

109

Agentic Index

Artificial Analysis
high

0.27

63

Rankings

Overall Rank

#46

Coding Rank

#35

About Gemini 3.5 Flash

Google's high-throughput, ultra-low-latency model debuted at Google I/O on May 19, 2026. Built directly for general availability scaling, it anchors Google's 24/7 proactive agentic infrastructure, dealing seamlessly with multimodal inputs and sprawling background tool operations.

Technical Specifications

Architecture specifications are undisclosed for proprietary models.

Attention

Attention Structure

Multi-Head Attention

Attention Heads

Undisclosed

Key-Value Heads

Undisclosed

Attention Head Dimension

Undisclosed

Position Embedding

Absolute Position Embedding

RoPE Theta

Undisclosed

Sliding Window Attention

Undisclosed

Sliding Window Size

Undisclosed

Sliding Window Ratio

Undisclosed

Linear Attention

Undisclosed

Linear Attention Ratio

Undisclosed

Normalization

Undisclosed

Activation Function

Undisclosed

Dimensions

Hidden Dimension Size

Undisclosed

Number of Layers

Undisclosed

FFN Intermediate Size (Dense)

Undisclosed

Multi-Token Prediction Heads

Undisclosed

Tokenizer

Vocabulary Size

Undisclosed

About Gemini 3.5

The Gemini 3.5 generation represents Google's transition into highly proactive, autonomous agent ecosystems. Skipping standard preview staging to target instant production scalability, it features highly optimized structural modalities tailored for multi-tool execution pipelines, persistent background agent actions, and sub-second core text response latency.


Other Gemini 3.5 Models
Gemini 3.5 Flash: Model Specifications and Details