ApX logoApX logo

Inkling

Active Parameters

975B

Context Length

1.05M

Modality

Multimodal

Architecture

Mixture of Experts (MoE)

License

Apache 2.0

Release Date

15 Jul 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $1.00 · Output: $4.05

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

1908.65 GB VRAM

Consumer

124x RTX 4090

24GB VRAM

Datacenter

31x NVIDIA A100

80GB VRAM

Apple Silicon

27x Apple M3 Max

128GB VRAM

1,048,576 tokens

2185.58 GB VRAM

Consumer

146x RTX 4090

24GB VRAM

Datacenter

36x NVIDIA A100

80GB VRAM

Apple Silicon

32x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RelativeHidden: 6.1k · Context: 1.05M · Vocab: 201kx 66 layersRMSNormPre-AttentionMulti-Head Attention64Q / 8KV heads · SW: 512Head dim: 128+RMSNormPre-FFNSparse MoE FFN (6/256 experts)ActivationIntermediate: 3.1k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#134

BenchmarkScoreRank
max

0.88

36

Agentic Coding

LiveBench Agentic
max

0.49

37

max

0.73

39

max

0.72

46

LiveBench Average

LiveBench Average
max

0.72

47

max

0.78

48

max

0.71

49

Agent Arena

Agent Arena

-10.9

53

General Text

Text Arena

1442

72

Web Development

WebDev Arena

1413

73

Agentic Index

Artificial Analysis
max

0.23

78

max

0.52

84

Intelligence Index

Artificial Analysis
max

0.25

120

Rankings

Overall Rank

#134

Coding Rank

#105

About Inkling

Inkling is a general-purpose open-weights 975B Mixture-of-Experts (MoE) multimodal model developed by Thinking Machine Labs. It features a 1M token context window, relative position embeddings for superior extrapolation, and native support for text, image, and audio inputs.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

64

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

Relative Position Embedding

RoPE Theta

-

Sliding Window Attention

Yes

Sliding Window Size

512

Sliding Window Ratio

83.3%

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

-

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

6,144

Number of Layers

66

FFN Intermediate Size (Dense)

24,576

Multi-Token Prediction Heads

8

Tokenizer

Vocabulary Size

201,024

Mixture of Experts

Total Expert Parameters

41.0B

Number of Experts

256

Active Experts

6

Shared Experts

2

FFN Intermediate Size (per Expert)

3,072

Dense Layers Before MoE

-

About Inkling

Inkling is a family of open-weights multimodal Mixture-of-Experts (MoE) models developed by Thinking Machine Labs.


Other Inkling Models