Active Parameters
975B
Context Length
1.05M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Apache 2.0
Release Date
15 Jul 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $1.00 · Output: $4.05
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
124x RTX 4090
24GB VRAM
Datacenter
31x NVIDIA A100
80GB VRAM
Apple Silicon
27x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
146x RTX 4090
24GB VRAM
Datacenter
36x NVIDIA A100
80GB VRAM
Apple Silicon
32x Apple M3 Max
128GB VRAM
Rank
#134
| Benchmark | Score | Rank |
|---|---|---|
Mathematics | max 0.88 | 36 |
Agentic Coding | max 0.49 | 37 |
Data Analysis | max 0.73 | 39 |
General | max 0.72 | 46 |
LiveBench Average | max 0.72 | 47 |
Reasoning | max 0.78 | 48 |
Coding | max 0.71 | 49 |
Agent Arena | -10.9 | 53 |
General Text | 1442 | 72 |
Web Development | 1413 | 73 |
Agentic Index | max 0.23 | 78 |
Coding Index | max 0.52 | 84 |
Intelligence Index | max 0.25 | 120 |
Overall Rank
#134
Coding Rank
#105
Inkling is a general-purpose open-weights 975B Mixture-of-Experts (MoE) multimodal model developed by Thinking Machine Labs. It features a 1M token context window, relative position embeddings for superior extrapolation, and native support for text, image, and audio inputs.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
64
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
Relative Position Embedding
RoPE Theta
-
Sliding Window Attention
Yes
Sliding Window Size
512
Sliding Window Ratio
83.3%
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
-
Dimensions
Auxiliary Parameters
-
Hidden Dimension Size
6,144
Number of Layers
66
FFN Intermediate Size (Dense)
24,576
Multi-Token Prediction Heads
8
Tokenizer
Vocabulary Size
201,024
Mixture of Experts
Total Expert Parameters
41.0B
Number of Experts
256
Active Experts
6
Shared Experts
2
FFN Intermediate Size (per Expert)
3,072
Dense Layers Before MoE
-
Inkling is a family of open-weights multimodal Mixture-of-Experts (MoE) models developed by Thinking Machine Labs.
Assistant
Online