ApX logoApX logo

Filters

Architecture

Source type

Status

Accuracy Scorecard

Comparing predictions against verified ground-truth configurations.

Open calculator

Calibration overview

Evaluation SuiteIn-Band Pass RateMedian ErrorGMFESystemic BiasWorst Case
VRAM Usage99.1%0.9%1.05x

Centered

E46 (1.93x)
Decode Speed94.9%13.9%1.20x

Underpredicts

T9 (0.30x)
Training Speed88.0%23.9%1.30x

Underpredicts

T23 (1.96x)

Memory predictions are highly deterministic and very accurate, while speed estimates are more sensitive to runtime and hardware configurations. Misses are largely due to unspecified parameter data in the source benchmarks.

Showing 114 of 114 cases

IDScenario & ConfigMeasuredPredictedRatioStatus

A1

Llama-3-8B (8.03B) fp16 weights = safetensors size16.1 GB16.1 GB

-0.1%

(1.00x)

Pass

A2

Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size19.9 GB19.5 GB

-1.7%

(0.98x)

Pass

A3

Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size26.9 GB27.2 GB

+1.1%

(1.01x)

Pass

A4

Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size34.8 GB33.3 GB

-4.4%

(0.96x)

Pass

A5

Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size47.4 GB45.3 GB

-4.3%

(0.96x)

Pass

A6

Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once26.4 GB27.6 GB

+4.2%

(1.04x)

Pass

A7

Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size14.5 GB14.5 GB

+0.0%

(1.00x)

Pass

A8

Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab)6.4 GB6.4 GB

-0.2%

(1.00x)

Pass

A9

Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit)4.7 GB4.8 GB

+1.7%

(1.02x)

Pass

A10

Llama-2-7B (6.74B) fp16 weights = safetensors size13.5 GB13.5 GB

+0.0%

(1.00x)

Pass

B1

Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128)1.1 GB1.1 GB

+0.0%

(1.00x)

Pass

B2

Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128)4.3 GB4.3 GB

+0.0%

(1.00x)

Pass

B3

Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128)1.3 GB1.3 GB

+0.0%

(1.00x)

Pass

B4

DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer)2.3 GB2.3 GB

+0.0%

(1.00x)

Pass

B4b

Kimi-K3 hybrid MLA + Linear Attention KV cache, 1M ctx, fp16 (93L, 576 cached dims/MLA layer, 74.2% linear reduction)29.0 GB29.0 GB

+0.0%

(1.00x)

Pass

B5

Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128)4.3 GB4.3 GB

+0.0%

(1.00x)

Pass

B6

Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128)2.7 GB2.7 GB

+0.0%

(1.00x)

Pass

B7

Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128)2.1 GB2.1 GB

+0.0%

(1.00x)

Pass

B8

Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128)1.9 GB1.9 GB

+0.0%

(1.00x)

Pass

B9

Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress17.2 GB17.2 GB

+0.0%

(1.00x)

Pass

D1

65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB)44.0 GB42.4 GB

-3.6%

(0.96x)

Pass

D2

Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth)8.5 GB8.8 GB

+3.3%

(1.03x)

Pass

D3

Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB)48.0 GB46.2 GB

-3.8%

(0.96x)

Pass

D4

7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations)112.0 GB130.5 GB

+16.5%

(1.17x)

Pass

D5

70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading70.0 GB45.3 GB

-35.2%

(0.65x)

Pass

D7

Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100)21.3 GB22.1 GB

+3.7%

(1.04x)

Pass

E1

End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded)17.5 GB17.5 GB

-0.2%

(1.00x)

Pass

E2

End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB)24.0 GB24.8 GB

+3.4%

(1.03x)

Pass

A11

Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight405.0 GB406.2 GB

+0.3%

(1.00x)

Pass

A12

Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight810.0 GB810.0 GB

+0.0%

(1.00x)

Pass

B10

Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress33.8 GB33.8 GB

+0.0%

(1.00x)

Pass

B11

DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims)10.3 GB10.3 GB

+0.0%

(1.00x)

Pass

B12

DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims)4.7 GB4.7 GB

+0.0%

(1.00x)

Pass

B18

DeepSeek-V4.1-Flash KV cache, 1M ctx, fp16 (40L CED, CSA2 cross-layer reuse)3.6 GB3.6 GB

+0.0%

(1.00x)

Pass

B19

DeepSeek-V4.1-Flash KV cache, 1M ctx, nvfp4 (40L CED, CSA2 890 B/tok)0.9 GB0.9 GB

+0.0%

(1.00x)

Pass

D8

Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload)54.7 GB43.9 GB

-19.7%

(0.80x)

Pass

E3

End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM)486.8 GB445.6 GB

-8.5%

(0.92x)

Pass

E4

End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM)946.2 GB856.9 GB

-9.4%

(0.91x)

Pass

D9

Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100)14.2 GB14.0 GB

-1.1%

(0.99x)

Pass

D10

Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU)36.7 GB32.5 GB

-11.3%

(0.89x)

Pass

D11

Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB)12.5 GB13.3 GB

+6.2%

(1.06x)

Pass

D12

Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling21.9 GB17.3 GB

-21.1%

(0.79x)

Pass

D13

Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB)17.4 GB20.9 GB

+20.1%

(1.20x)

Pass

D14

Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB)7.9 GB10.5 GB

+32.9%

(1.33x)

Pass

D15

Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU)29.6 GB25.4 GB

-14.2%

(0.86x)

Pass

D16

Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU)44.8 GB49.4 GB

+10.3%

(1.10x)

Pass

D17

Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB)24.9 GB23.9 GB

-4.2%

(0.96x)

Pass

D18

33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B21.0 GB20.9 GB

-0.5%

(0.99x)

Pass

E5

PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx8.3 GB7.0 GB

-16.0%

(0.84x)

Pass

E6

PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx31.6 GB28.3 GB

-10.6%

(0.89x)

Pass

E7

PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx475.2 GB342.8 GB

-27.9%

(0.72x)

Pass

E9

PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx215.4 GB225.1 GB

+4.5%

(1.04x)

Pass

A13

Qwen3-Omni-30B-A3B (33.0B total) BF16 weights = safetensors size66.0 GB66.0 GB

+0.0%

(1.00x)

Pass

A14

Qwen3-Omni-30B-A3B (33.0B total, 3B aux) W4A16 weights = AutoRound int4 size25.0 GB23.8 GB

-4.8%

(0.95x)

Pass

E10

End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE)2100.0 GB2073.6 GB

-1.3%

(0.99x)

Pass

B13

Llama-3.1-70B KV Cache with batch size 3285.9 GB85.9 GB

+0.0%

(1.00x)

Pass

D19

DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes)76.5 GB81.0 GB

+5.9%

(1.06x)

Pass

A15

Inkling-Small (276.0B / 12.0B active) BF16 weights = base safetensors size552.0 GB552.0 GB

+0.0%

(1.00x)

Pass

A16

Inkling-Small (276.0B / 12.0B active) NVFP4 weights140.0 GB140.5 GB

+0.3%

(1.00x)

Pass

A17

Nemotron 3.5 Lightning (30.0B / 3.0B active) BF16 weights = safetensors size60.0 GB60.0 GB

+0.0%

(1.00x)

Pass

A18

Nemotron 3.5 Lightning (30.0B / 3.0B active) NVFP4 weights16.0 GB16.1 GB

+0.4%

(1.00x)

Pass

A19

Qwen3.8-2.4T-A95B (2400.0B / 95.0B active) NVFP4 weights1200.0 GB1206.1 GB

+0.5%

(1.01x)

Pass

B14

Inkling-Small KV cache, 8k ctx, fp16 (42L, 8 KV heads, head 128)1.4 GB1.4 GB

+0.0%

(1.00x)

Pass

B15

Nemotron 3.5 Lightning KV cache, 8k ctx, fp16 (52L, hidden 2688, linear attention ratio 0.75)1.1 GB1.1 GB

+0.0%

(1.00x)

Pass

A20

Muse-Glimmer-30B (29.6B) bf16 weights = safetensors size59.2 GB59.2 GB

+0.0%

(1.00x)

Pass

B16

Muse-Glimmer-30B KV cache, 8k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048)0.2 GB0.2 GB

+0.0%

(1.00x)

Pass

B17

Muse-Glimmer-30B KV cache, 128k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048)1.8 GB1.8 GB

+0.0%

(1.00x)

Pass

E12

End-to-end: Muse-Glimmer-30B Q4_K_M inference, 8k ctx (claimed ~17-24GB on RTX 4090/5090)19.5 GB22.2 GB

+14.1%

(1.14x)

Pass

E13

Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 1 (vLLM memory profiling: ~140GB weights + ~42GB KV + ~19GB framework/NCCL)203.0 GB203.6 GB

+0.3%

(1.00x)

Pass

E14

Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 4 (vLLM memory profiling with chunked prefill: ~205GB total across cluster)205.5 GB205.8 GB

+0.1%

(1.00x)

Pass

A21

Llama-3.2-11B-Vision (10.67B total, 2.64B aux) BF16 weights = safetensors size21.3 GB21.3 GB

+0.0%

(1.00x)

Pass

A22

Qwen2-VL-7B-Instruct (8.29B untied safetensors) BF16 weights = safetensors size16.6 GB16.6 GB

+0.0%

(1.00x)

Pass

A23

Qwen2-VL-7B-Instruct-GPTQ-Int4 weights = GPTQ Int4 safetensors size6.9 GB6.8 GB

-1.9%

(0.98x)

Pass

A24

Qwen2.5-3B-Instruct (3.09B) BF16 weights = published safetensors size6.2 GB6.2 GB

+0.2%

(1.00x)

Pass

A25

Mistral-7B-Instruct-v0.3 (7.25B) BF16 weights = published safetensors size14.5 GB14.5 GB

+0.0%

(1.00x)

Pass

A26

Qwen2.5-0.5B-Instruct (0.49B) BF16 weights = published single safetensors size1.0 GB1.0 GB

-1.0%

(0.99x)

Pass

A27

Qwen3.5-0.8B (0.87B total with 0.10B visual encoder) BF16 weights = published safetensors size1.8 GB1.7 GB

-0.6%

(0.99x)

Pass

A28

Gemma 4 E2B (5.12B total with 0.49B vision+audio) BF16 weights = published safetensors size10.3 GB10.2 GB

-0.1%

(1.00x)

Pass

B20

Qwen2.5-3B KV cache, 8k ctx, fp16 (36L, 2 KV heads, head 128)0.3 GB0.3 GB

+7.1%

(1.07x)

Pass

E15

Qwen2.5-3B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 8.01GB)8.0 GB8.1 GB

+0.9%

(1.01x)

Pass

E16

Qwen2.5-3B-Instruct FP16 raw unchunked prefill, RTX 4090 (24GB), 8k ctx, batch 4 -> OOMInsufficientInsufficient-
Pass

E17

DeepSeek-V2-Lite (16B MoE total / 2.4B active) FP16 weights loading on RTX 4090 (24GB) -> OOMInsufficientInsufficient-
Pass

E18

Qwen2-VL-2B (2.21B total with auxiliary visual encoder) BF16 empirical weights on RTX 40904.4 GB4.4 GB

+0.0%

(1.00x)

Pass

E19

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 15.41GB)15.4 GB16.5 GB

+6.9%

(1.07x)

Pass

E20

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 4 (measured peak: 18.59GB)18.6 GB19.3 GB

+3.9%

(1.04x)

Pass

E21

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 8 (measured peak: 23.05GB)23.1 GB23.6 GB

+2.6%

(1.03x)

Pass

E22

Qwen2.5-7B-Instruct FP16 empirical model weights on RTX 409015.2 GB15.2 GB

+0.1%

(1.00x)

Pass

E23

Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 1k ctx, batch 1 (measured peak: 18.40GB)18.4 GB19.9 GB

+8.3%

(1.08x)

Pass

E24

Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 19.87GB)19.9 GB20.3 GB

+2.3%

(1.02x)

Pass

E25

Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 1 (measured peak: 17.19GB)17.2 GB17.7 GB

+3.2%

(1.03x)

Pass

E26

Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 8 (measured peak: 32.73GB)32.7 GB28.2 GB

-13.8%

(0.86x)

Pass

E27

Llama-3.1-8B-Instruct QLoRA r=16 training, 1x RTX 4090 (24GB), bs=2 seq=2048 (measured peak: 19.71GB)19.7 GB15.0 GB

-23.7%

(0.76x)

Pass

E28

Qwen2.5-3B-Instruct full FT, 1x A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 33.59GB)33.6 GB32.6 GB

-3.0%

(0.97x)

Pass

E29

DeepSeek-R1-Distill-Qwen-14B FP16 empirical model weights, A100-SXM4 (80GB)29.5 GB29.5 GB

+0.0%

(1.00x)

Pass

E30

DeepSeek-R1-Distill-Qwen-14B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 30.30GB)30.3 GB31.1 GB

+2.8%

(1.03x)

Pass

E31

DeepSeek-R1-Distill-Qwen-7B FP16 empirical model weights, RTX 4090 (24GB)15.2 GB15.2 GB

+0.1%

(1.00x)

Pass

E32

DeepSeek-R1-Distill-Qwen-7B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 16.23GB)16.2 GB16.5 GB

+1.5%

(1.02x)

Pass

E33

DeepSeek-R1-Distill-Qwen-32B FP16 empirical model weights, A100-SXM4 (80GB)65.5 GB65.5 GB

-0.0%

(1.00x)

Pass

E34

DeepSeek-R1-Distill-Qwen-32B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 66.65GB)66.7 GB67.3 GB

+1.1%

(1.01x)

Pass

E35

Qwen2.5-3B-Instruct FP16 inference, H200 (141GB), 2k ctx, batch 1 (measured peak: 6.76GB)6.8 GB7.3 GB

+8.3%

(1.08x)

Pass

E36

Qwen3.8-27B FP16 empirical model weights, A100-SXM4 (80GB)53.8 GB53.8 GB

+0.0%

(1.00x)

Pass

E37

Qwen3.8-27B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 54.67GB)54.7 GB55.3 GB

+1.2%

(1.01x)

Pass

E38

Ministral-3-8B-Instruct-2512 FP16 empirical model weights, A100-SXM4 (80GB)17.8 GB17.8 GB

+0.0%

(1.00x)

Pass

E39

Ministral-3-8B-Instruct-2512 FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 20.38GB)20.4 GB19.3 GB

-5.5%

(0.94x)

Pass

E40

Granite-4.2-8B FP16 empirical model weights, RTX 4090 (24GB)17.6 GB17.6 GB

+0.0%

(1.00x)

Pass

E41

Granite-4.2-8B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 18.89GB)18.9 GB19.1 GB

+0.9%

(1.01x)

Pass

E42

Gemma-4-E2B FP16 empirical model weights, RTX 4090 (24GB)10.2 GB10.2 GB

-0.1%

(1.00x)

Pass

E43

Gemma-4-E2B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 12.81GB)12.8 GB13.1 GB

+2.6%

(1.03x)

Pass

E44

Gemma-4-12B FP16 empirical model weights, L40S (48GB)23.9 GB23.9 GB

-0.1%

(1.00x)

Pass

E45

Gemma-4-12B FP16 inference, L40S (48GB), 2k ctx, batch 1 (measured peak: 26.62GB)26.6 GB26.6 GB

-0.2%

(1.00x)

Pass

E46

Ministral-3-3B Full FT 32-bit AdamW, A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 36.95GB)37.0 GB71.5 GB

+93.4%

(1.93x)

Miss

E47

DeepSeek-R1-Distill-Qwen-7B LoRA 32-bit AdamW training, RTX 4090 (24GB) -> OOMInsufficientInsufficient-
Pass

E48

Ministral-3-8B-Instruct-2512 LoRA 8-bit AdamW training, RTX 4090 (24GB) -> OOMInsufficientInsufficient-
Pass

E49

Qwen3.8-Flash-Next (176.9B) FP16 inference, 2x H200 (141GB) -> OOM (weights exceed cluster capacity)InsufficientInsufficient-
Pass