Active Parameters
196.81B
Context Length
256K
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Apache 2.0
Release Date
11 Feb 2026
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
22x RTX 4090
24GB VRAM
Datacenter
6x NVIDIA A100
80GB VRAM
Apple Silicon
5x Apple M3 Max
128GB VRAM
256,000 tokens
Consumer
24x RTX 4090
24GB VRAM
Datacenter
7x NVIDIA A100
80GB VRAM
Apple Silicon
5x Apple M3 Max
128GB VRAM
Rank
#68
| Benchmark | Score | Rank |
|---|---|---|
General Text | 1394 | 100 |
Overall Rank
#68
Coding Rank
-
Step 3.5 Flash is StepFun's most capable open-source foundation model. Engineered on a sparse Mixture of Experts architecture with 196B total parameters but only 11B active per token, it delivers frontier reasoning and agentic capabilities with exceptional efficiency. Features 256K context window, supports text and image input, and achieves 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0. Optimized for local deployment on consumer hardware including Mac Studio M4 Max and high-end GPUs. Powered by 3-way Multi-Token Prediction (MTP-3) for 100-350 tok/s generation throughput.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
64
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
Yes
Sliding Window Size
512
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
GELU
Dimensions
Hidden Dimension Size
4,096
Number of Layers
45
FFN Intermediate Size (Dense)
1,280
Multi-Token Prediction Heads
3
Tokenizer
Vocabulary Size
128,896
Mixture of Experts
Total Expert Parameters
11.0B
Number of Experts
288
Active Experts
8
Shared Experts
-
FFN Intermediate Size (per Expert)
1,280
Dense Layers Before MoE
-
Step 3.5 is StepFun's flagship frontier reasoning model family. Built on sparse Mixture-of-Experts (MoE) architecture, Step 3.5 models deliver frontier-level intelligence for agentic, reasoning, and coding tasks. The Flash variant selectively activates only 11B of its 196B parameters per token, achieving the reasoning depth of top-tier proprietary models while maintaining exceptional efficiency. Features 256K context window, native function calling, and Multi-Token Prediction for high-throughput inference. Released under Apache 2.0 license.
APX AI
Online