Active Parameters
276B
Context Length
1.05M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Apache 2.0
Release Date
30 Jul 2026
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
31x RTX 4090
24GB VRAM
Datacenter
9x NVIDIA A100
80GB VRAM
Apple Silicon
7x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
43x RTX 4090
24GB VRAM
Datacenter
12x NVIDIA A100
80GB VRAM
Apple Silicon
9x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Inkling-Small available.
Overall Rank
-
Coding Rank
-
Inkling-Small is an open-weights 276B Mixture-of-Experts (MoE) multimodal foundation model with 12B active parameters per token developed by Thinking Machine Labs. Trained on 45T multimodal tokens, it supports text, image, and native audio inputs with up to a 1M token context window, variable thinking effort, and sliding-window attention. It achieves near-parity and on certain benchmarks outperforms the 975B flagship Inkling at substantially lower inference latency and footprint.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
32
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
Relative Position Embedding
RoPE Theta
-
Sliding Window Attention
Yes
Sliding Window Size
512
Sliding Window Ratio
83.3%
Linear Attention
No
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
-
Dimensions
Hidden Dimension Size
4,096
Number of Layers
42
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
201,024
Mixture of Experts
Total Expert Parameters
12.0B
Number of Experts
128
Active Experts
6
Shared Experts
2
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
-
Inkling is a family of open-weights multimodal Mixture-of-Experts (MoE) models developed by Thinking Machine Labs.
APX AI
Online