ApX logoApX logo

Filters

Architecture

Source type

Status

Accuracy Scorecard

Comparing predictions against verified ground-truth configurations.

Open calculator

Calibration overview

Evaluation SuiteIn-Band Pass RateMedian ErrorGMFESystemic BiasWorst Case
VRAM Usage92.6%1.8%1.05x

Centered

E46 (1.90x)
Decode Speed88.1%8.4%1.13x

Slight Underpredict

T9 (0.30x)
Training Speed90.4%6.2%1.11x

Slight Underpredict

T22 (0.55x)

Memory predictions are highly deterministic and very accurate, while speed estimates are more sensitive to runtime and hardware configurations. Misses are largely due to unspecified parameter data in the source benchmarks.

Showing 270 of 270 cases

IDScenario & ConfigMeasuredPredictedRatioStatus

A1

Llama-3-8B (8.03B) fp16 weights = safetensors size15.0 GB15.0 GB

-0.1%

(1.00x)

Pass

A2

Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size18.5 GB18.2 GB

-1.7%

(0.98x)

Pass

A3

Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size25.1 GB25.3 GB

+1.1%

(1.01x)

Pass

A4

Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size32.4 GB31.0 GB

-4.4%

(0.96x)

Pass

A5

Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size44.1 GB42.2 GB

-4.3%

(0.96x)

Pass

A6

Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once24.6 GB25.7 GB

+4.2%

(1.04x)

Pass

A7

Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size13.5 GB13.5 GB

+0.0%

(1.00x)

Pass

A8

Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab)6.0 GB6.0 GB

-0.2%

(1.00x)

Pass

A9

Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit)4.4 GB4.4 GB

+1.8%

(1.02x)

Pass

A10

Llama-2-7B (6.74B) fp16 weights = safetensors size12.6 GB12.6 GB

+0.0%

(1.00x)

Pass

B1

Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128)1.0 GB1.0 GB

+0.0%

(1.00x)

Pass

B2

Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128)4.0 GB4.0 GB

+0.0%

(1.00x)

Pass

B3

Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128)1.3 GB1.3 GB

+0.0%

(1.00x)

Pass

B4

DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer)2.1 GB2.1 GB

+0.0%

(1.00x)

Pass

B4b

Kimi-K3 hybrid MLA + Linear Attention KV cache, 1M ctx, fp16 (93L, 576 cached dims/MLA layer, 74.2% linear reduction)27.0 GB27.0 GB

+0.0%

(1.00x)

Pass

B5

Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128)4.0 GB4.0 GB

+0.0%

(1.00x)

Pass

B6

Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128)2.5 GB2.5 GB

+0.0%

(1.00x)

Pass

B7

Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128)2.0 GB2.0 GB

+0.0%

(1.00x)

Pass

B8

Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128)1.8 GB1.8 GB

+0.0%

(1.00x)

Pass

B9

Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress16.0 GB16.0 GB

+0.0%

(1.00x)

Pass

D1

65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB)44.0 GB42.9 GB

-2.6%

(0.97x)

Pass

D2

Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth)8.5 GB7.4 GB

-12.8%

(0.87x)

Miss

D3

Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB)48.0 GB42.7 GB

-11.0%

(0.89x)

Miss

D4

7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations)104.3 GB122.3 GB

+17.3%

(1.17x)

Pass

D5

70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading70.0 GB47.0 GB

-32.8%

(0.67x)

Pass

D7

Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100)21.3 GB21.2 GB

-0.5%

(1.00x)

Pass

E1

End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded)17.5 GB16.3 GB

-6.6%

(0.93x)

Miss

E2

End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB)24.0 GB23.2 GB

-3.3%

(0.97x)

Pass

A11

Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight377.2 GB378.3 GB

+0.3%

(1.00x)

Pass

A12

Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight754.4 GB754.4 GB

+0.0%

(1.00x)

Pass

B10

Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress31.5 GB31.5 GB

+0.0%

(1.00x)

Pass

B11

DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims)9.6 GB9.6 GB

+0.0%

(1.00x)

Pass

B12

DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims)4.3 GB4.3 GB

+0.0%

(1.00x)

Pass

B18

DeepSeek-V4.1-Flash KV cache, 1M ctx, fp16 (40L CED, CSA2 cross-layer reuse)3.4 GB3.4 GB

+0.0%

(1.00x)

Pass

B19

DeepSeek-V4.1-Flash KV cache, 1M ctx, nvfp4 (40L CED, CSA2 890 B/tok)0.9 GB0.9 GB

+0.0%

(1.00x)

Pass

D8

Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload)54.7 GB45.7 GB

-16.4%

(0.84x)

Pass

E3

End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM)486.8 GB417.2 GB

-14.3%

(0.86x)

Miss

E4

End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM)881.2 GB800.3 GB

-9.2%

(0.91x)

Pass

D9

Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100)14.2 GB13.8 GB

-2.7%

(0.97x)

Pass

D10

Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU)36.7 GB31.0 GB

-15.4%

(0.85x)

Pass

D11

Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB)12.5 GB13.1 GB

+4.4%

(1.04x)

Pass

D12

Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling21.9 GB16.9 GB

-23.1%

(0.77x)

Pass

D13

Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB)17.4 GB20.9 GB

+20.3%

(1.20x)

Pass

D14

Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB)7.9 GB10.3 GB

+30.0%

(1.30x)

Pass

D15

Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU)29.6 GB24.9 GB

-15.7%

(0.84x)

Pass

D16

Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU)44.8 GB45.3 GB

+1.1%

(1.01x)

Pass

D17

Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB)24.9 GB22.9 GB

-7.9%

(0.92x)

Pass

D18

33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B21.0 GB21.7 GB

+3.3%

(1.03x)

Pass

E5

PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx8.3 GB6.6 GB

-20.8%

(0.79x)

Miss

E6

PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx31.6 GB26.4 GB

-16.5%

(0.84x)

Miss

E7

PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx475.2 GB319.4 GB

-32.8%

(0.67x)

Miss

E9

PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx215.4 GB209.8 GB

-2.6%

(0.97x)

Pass

A13

Qwen3-Omni-30B-A3B (33.0B total) BF16 weights = safetensors size61.5 GB61.5 GB

+0.0%

(1.00x)

Pass

A14

Qwen3-Omni-30B-A3B (33.0B total, 3B aux) W4A16 weights = AutoRound int4 size23.3 GB22.2 GB

-4.8%

(0.95x)

Pass

E10

End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE)2100.0 GB2073.6 GB

-1.3%

(0.99x)

Pass

B13

Llama-3.1-70B KV Cache with batch size 3280.0 GB80.0 GB

+0.0%

(1.00x)

Pass

D19

DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes)76.5 GB80.3 GB

+4.9%

(1.05x)

Pass

A15

Inkling-Small (276.0B / 12.0B active) BF16 weights = base safetensors size514.1 GB514.1 GB

+0.0%

(1.00x)

Pass

A16

Inkling-Small (276.0B / 12.0B active) NVFP4 weights130.4 GB130.8 GB

+0.3%

(1.00x)

Pass

A17

Nemotron 3.5 Lightning (30.0B / 3.0B active) BF16 weights = safetensors size55.9 GB55.9 GB

+0.0%

(1.00x)

Pass

A18

Nemotron 3.5 Lightning (30.0B / 3.0B active) NVFP4 weights14.9 GB14.9 GB

+0.3%

(1.00x)

Pass

A19

Qwen3.8-2.4T-A95B (2400.0B / 95.0B active) NVFP4 weights1117.6 GB1123.3 GB

+0.5%

(1.01x)

Pass

B14

Inkling-Small KV cache, 8k ctx, fp16 (42L, 8 KV heads, head 128)1.3 GB1.3 GB

+0.0%

(1.00x)

Pass

B15

Nemotron 3.5 Lightning KV cache, 8k ctx, fp16 (52L, hidden 2688, linear attention ratio 0.75)1.1 GB1.1 GB

+0.0%

(1.00x)

Pass

A20

Muse-Glimmer-30B (29.6B) bf16 weights = safetensors size55.1 GB55.1 GB

+0.0%

(1.00x)

Pass

B16

Muse-Glimmer-30B KV cache, 8k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048)0.2 GB0.2 GB

+0.0%

(1.00x)

Pass

B17

Muse-Glimmer-30B KV cache, 128k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048)1.7 GB1.7 GB

+0.0%

(1.00x)

Pass

E12

End-to-end: Muse-Glimmer-30B Q4_K_M inference, 8k ctx (claimed ~17-24GB on RTX 4090/5090)19.5 GB20.8 GB

+6.6%

(1.07x)

Pass

E13

Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 1 (vLLM memory profiling: ~140GB weights + ~42GB KV + ~19GB framework/NCCL)203.0 GB205.5 GB

+1.2%

(1.01x)

Pass

E14

Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 4 (vLLM memory profiling with chunked prefill: ~205GB total across cluster)205.5 GB205.5 GB

-0.0%

(1.00x)

Pass

A21

Llama-3.2-11B-Vision (10.67B total, 2.64B aux) BF16 weights = safetensors size19.9 GB19.9 GB

+0.0%

(1.00x)

Pass

A22

Qwen2-VL-7B-Instruct (8.29B untied safetensors) BF16 weights = safetensors size15.4 GB15.4 GB

+0.0%

(1.00x)

Pass

A23

Qwen2-VL-7B-Instruct-GPTQ-Int4 weights = GPTQ Int4 safetensors size6.5 GB6.3 GB

-1.9%

(0.98x)

Pass

A24

Qwen2.5-3B-Instruct (3.09B) BF16 weights = published safetensors size5.8 GB5.8 GB

+0.2%

(1.00x)

Pass

A25

Mistral-7B-Instruct-v0.3 (7.25B) BF16 weights = published safetensors size13.5 GB13.5 GB

+0.0%

(1.00x)

Pass

A26

Qwen2.5-0.5B-Instruct (0.49B) BF16 weights = published single safetensors size0.9 GB0.9 GB

-1.1%

(0.99x)

Pass

A27

Qwen3.5-0.8B (0.87B total with 0.10B visual encoder) BF16 weights = published safetensors size1.6 GB1.6 GB

-0.6%

(0.99x)

Pass

A28

Gemma 4 E2B (5.12B total with 0.49B vision+audio) BF16 weights = published safetensors size9.6 GB9.5 GB

-0.1%

(1.00x)

Pass

B20

Qwen2.5-3B KV cache, 8k ctx, fp16 (36L, 2 KV heads, head 128)0.3 GB0.3 GB

+0.0%

(1.00x)

Pass

E15

Qwen2.5-3B-Instruct FP16 raw forward (all-position logits), RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 6.85GB)6.8 GB7.3 GB

+7.0%

(1.07x)

Pass

E16

Qwen2.5-3B-Instruct FP16 raw forward (all-position logits), RTX 4090 (24GB), 8k ctx, batch 4 (measured peak: 19.773GB; fits - the earlier OOM was the harness holding two logits tensors)19.8 GB19.7 GB

-0.6%

(0.99x)

Pass

E17

DeepSeek-V2-Lite (16B MoE total / 2.4B active) FP16 weights loading on RTX 4090 (24GB) -> OOMInsufficientInsufficient-
Pass

E18

Qwen2-VL-2B (2.21B total with auxiliary visual encoder) BF16 empirical weights on RTX 40904.1 GB4.1 GB

+0.0%

(1.00x)

Pass

E19

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 15.41GB)15.4 GB15.9 GB

+3.2%

(1.03x)

Pass

E20

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 4 (measured peak: 18.59GB)18.6 GB18.7 GB

+0.4%

(1.00x)

Pass

E21

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 8 (measured peak: 23.05GB)23.1 GB22.8 GB

-1.0%

(0.99x)

Pass

E22

Qwen2.5-7B-Instruct FP16 empirical model weights on RTX 409014.2 GB14.2 GB

+0.1%

(1.00x)

Pass

E23

Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 1k ctx, batch 1 (measured peak: 18.40GB)18.4 GB19.3 GB

+5.0%

(1.05x)

Pass

E24

Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 19.87GB)19.9 GB19.7 GB

-0.9%

(0.99x)

Pass

E25

Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 1 (measured peak: 17.19GB)17.2 GB17.2 GB

+0.2%

(1.00x)

Pass

E26

Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 8 (measured peak: 32.73GB)32.7 GB30.1 GB

-8.1%

(0.92x)

Pass

E27

Llama-3.1-8B-Instruct QLoRA r=16 training, 1x RTX 4090 (24GB), bs=2 seq=2048 (measured peak: 19.71GB)19.7 GB19.7 GB

-0.2%

(1.00x)

Pass

E28

Qwen2.5-3B-Instruct full FT, 1x A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 33.59GB)33.6 GB32.1 GB

-4.3%

(0.96x)

Pass

E29

DeepSeek-R1-Distill-Qwen-14B FP16 empirical model weights, A100-SXM4 (80GB)27.5 GB27.5 GB

+0.0%

(1.00x)

Pass

E30

DeepSeek-R1-Distill-Qwen-14B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 30.30GB)30.3 GB30.5 GB

+0.7%

(1.01x)

Pass

E31

DeepSeek-R1-Distill-Qwen-7B FP16 empirical model weights, RTX 4090 (24GB)14.2 GB14.2 GB

+0.1%

(1.00x)

Pass

E32

DeepSeek-R1-Distill-Qwen-7B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 16.23GB)16.2 GB15.9 GB

-2.0%

(0.98x)

Pass

E33

DeepSeek-R1-Distill-Qwen-32B FP16 empirical model weights, A100-SXM4 (80GB)61.0 GB61.0 GB

-0.0%

(1.00x)

Pass

E34

DeepSeek-R1-Distill-Qwen-32B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 66.65GB)66.7 GB66.7 GB

+0.0%

(1.00x)

Pass

E35

Qwen2.5-3B-Instruct FP16 inference, H200 (141GB), 2k ctx, batch 1 (measured peak: 6.76GB)6.8 GB6.8 GB

+0.0%

(1.00x)

Pass

E36

Qwen3.8-27B FP16 empirical model weights, A100-SXM4 (80GB)50.1 GB50.1 GB

+0.0%

(1.00x)

Pass

E37

Qwen3.8-27B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 54.67GB)54.7 GB54.7 GB

-0.0%

(1.00x)

Pass

E38

Ministral-3-8B-Instruct-2512 FP16 empirical model weights, A100-SXM4 (80GB)16.6 GB16.6 GB

+0.0%

(1.00x)

Pass

E39

Ministral-3-8B-Instruct-2512 FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 20.38GB)20.4 GB18.7 GB

-8.4%

(0.92x)

Pass

E40

Granite-4.2-8B FP16 empirical model weights, RTX 4090 (24GB)16.4 GB16.4 GB

+0.0%

(1.00x)

Pass

E41

Granite-4.2-8B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 18.89GB)18.9 GB18.5 GB

-2.3%

(0.98x)

Pass

E42

Gemma-4-E2B FP16 empirical model weights, RTX 4090 (24GB)9.5 GB9.5 GB

-0.1%

(1.00x)

Pass

E43

Gemma-4-E2B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 12.81GB)12.8 GB12.4 GB

-2.8%

(0.97x)

Pass

E44

Gemma-4-12B FP16 empirical model weights, L40S (48GB)22.3 GB22.3 GB

-0.1%

(1.00x)

Pass

E45

Gemma-4-12B FP16 inference, L40S (48GB), 2k ctx, batch 1 (measured peak: 26.62GB)26.6 GB25.9 GB

-2.8%

(0.97x)

Pass

E46

Ministral-3-3B Full FT 32-bit AdamW, A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 36.95GB)37.0 GB70.1 GB

+89.7%

(1.90x)

Miss

E47

DeepSeek-R1-Distill-Qwen-7B LoRA 32-bit AdamW training, RTX 4090 (24GB) -> OOMInsufficientInsufficient-
Pass

E48

Ministral-3-8B-Instruct-2512 LoRA 8-bit AdamW training, RTX 4090 (24GB) -> OOMInsufficientInsufficient-
Pass

E49

Qwen3.8-Flash-Next (176.9B) FP16 inference, 2x H200 (141GB) -> OOM (weights exceed cluster capacity)InsufficientInsufficient-
Pass

E50

Qwen2.5-3B-Instruct FP16 prefill forward with last-token logits (as generate() does), RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 6.27GB)6.3 GB6.8 GB

+7.8%

(1.08x)

Pass

E51

Qwen2.5-3B-Instruct FP16 prefill forward with last-token logits (as generate() does), RTX 4090 (24GB), 8k ctx, batch 4 (measured peak: 10.5GB)10.5 GB10.4 GB

-1.0%

(0.99x)

Pass

E52

Gemma-4-12B BF16 inference, 2x B300 SXM6 (288GB), 2k ctx, batch 1 (measured physical peak: 28.72GB)28.7 GB26.3 GB

-8.5%

(0.92x)

Pass

E53

Devstral-2-123B-Instruct FP8 inference, 2x B300 SXM6 (288GB), 2k ctx, batch 1 (measured physical peak: 128.15GB)128.2 GB139.8 GB

+9.1%

(1.09x)

Pass

E54

Qwen2.5-7B-Instruct FP16 eager inference, AMD Instinct MI350X (288GB), 2k ctx, batch 1 (measured physical peak: 16.03GB)16.0 GB15.9 GB

-0.9%

(0.99x)

Pass

VM1

Qwen2.5-32B-Instruct BF16, vLLM 0.30.0 defaults, 2x H100 SXM (80GB) TP=2: KV cache blocks after profiling (19691 x 16 tokens, gpu_memory_utilization 0.92)19691.0 GB19847.0 GB

+0.8%

(1.01x)

Pass

VM2

Qwen2.5-3B-Instruct BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (25334 x 16 tokens, gpu_memory_utilization 0.92)25334.0 GB25158.0 GB

-0.7%

(0.99x)

Pass

VM3

Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (5422 x 16 tokens, gpu_memory_utilization 0.92)5422.0 GB6225.0 GB

+14.8%

(1.15x)

Pass

VM4

DeepSeek-R1-Distill-Qwen-7B BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (6567 x 16 tokens, gpu_memory_utilization 0.92)6567.0 GB6225.0 GB

-5.2%

(0.95x)

Pass

VM5

Qwen2.5-14B-Instruct BF16, vLLM 0.30.0 defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (14195 x 16 tokens, gpu_memory_utilization 0.92)14195.0 GB14824.0 GB

+4.4%

(1.04x)

Pass

VM6

Qwen2.5-14B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (13630 x 16 tokens, gpu_memory_utilization 0.92)13630.0 GB14704.0 GB

+7.9%

(1.08x)

Pass

VM7

Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (64060 x 16 tokens, gpu_memory_utilization 0.92)64060.0 GB66226.0 GB

+3.4%

(1.03x)

Pass

VM8

Llama-3.1-8B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (27701 x 16 tokens, gpu_memory_utilization 0.92)27701.0 GB28554.0 GB

+3.1%

(1.03x)

Pass

FM1

Llama-3.1-8B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 14.07 GiB14.1 GB15.4 GB

+9.2%

(1.09x)

Pass

FM2

Llama-3.1-8B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOMInsufficientInsufficient-
Pass

FM3

Llama-3.1-8B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 15.06 GiB15.1 GB16.9 GB

+11.9%

(1.12x)

Miss

FM4

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 13.87 GiB13.9 GB14.2 GB

+2.5%

(1.02x)

Pass

FM5

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 21.11 GiB21.1 GB19.6 GB

-7.3%

(0.93x)

Pass

FM6

Qwen2.5-3B-Instruct LORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOMInsufficientInsufficient-
Pass

FM7

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 16.45 GiB16.4 GB16.4 GB

-0.4%

(1.00x)

Pass

FM8

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 22.82 GiB22.8 GB23.9 GB

+4.8%

(1.05x)

Pass

FM9

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs4 seq512: peak device memory 21.11 GiB21.1 GB19.6 GB

-7.3%

(0.93x)

Pass

FM10

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 19.91 GiB19.9 GB21.6 GB

+8.5%

(1.08x)

Pass

FM11

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 14.43 GiB14.4 GB15.1 GB

+4.8%

(1.05x)

Pass

FM12

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 22.69 GiB22.7 GB21.9 GB

-3.7%

(0.96x)

Pass

FM13

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOMInsufficientInsufficient-
Pass

FM14

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 15.7 GiB15.7 GB17.1 GB

+9.2%

(1.09x)

Pass

FM15

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs4 seq2048 -> OOMInsufficientInsufficient-
Pass

FM16

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 10.32 GiB10.3 GB10.7 GB

+3.5%

(1.03x)

Pass

FM17

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 12.88 GiB12.9 GB12.9 GB

-0.1%

(1.00x)

Pass

FM18

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 18.01 GiB18.0 GB17.8 GB

-1.3%

(0.99x)

Pass

FM19

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 18.01 GiB18.0 GB17.3 GB

-4.2%

(0.96x)

Pass

FM20

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 18.87 GiB18.9 GB19.5 GB

+3.3%

(1.03x)

Pass

FM21

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 21.83 GiB21.8 GB22.1 GB

+1.1%

(1.01x)

Pass

FM22

Qwen2.5-7B-Instruct LORA r=16, torchtune None defaults, 1x RTX 4090 (24GB), bs2 seq4096 -> OOMInsufficientInsufficient-
Pass

FM23

Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 16.46 GiB16.5 GB17.0 GB

+3.2%

(1.03x)

Pass

FM24

Llama-3.1-8B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 8.52 GiB8.5 GB8.6 GB

+0.6%

(1.01x)

Pass

FM25

Llama-3.1-8B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.55 GiB7.5 GB7.7 GB

+1.7%

(1.02x)

Pass

FM26

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 7.08 GiB7.1 GB7.3 GB

+3.7%

(1.04x)

Pass

FM27

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 7.94 GiB7.9 GB8.3 GB

+3.9%

(1.04x)

Pass

FM28

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.55 GiB7.5 GB7.5 GB

-0.4%

(1.00x)

Pass

FM29

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 7.94 GiB7.9 GB8.0 GB

+0.8%

(1.01x)

Pass

FM30

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 7.94 GiB7.9 GB7.9 GB

-0.9%

(0.99x)

Pass

FM31

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 16.26 GiB16.3 GB16.4 GB

+0.7%

(1.01x)

Pass

FM32

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 6.89 GiB6.9 GB7.2 GB

+3.9%

(1.04x)

Pass

FM33

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 8.01 GiB8.0 GB8.2 GB

+2.7%

(1.03x)

Pass

FM34

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.38 GiB7.4 GB7.4 GB

+0.1%

(1.00x)

Pass

FM35

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 8.01 GiB8.0 GB8.0 GB

-0.4%

(1.00x)

Pass

FM36

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 8.01 GiB8.0 GB7.8 GB

-2.0%

(0.98x)

Pass

FM37

Phi-3-mini-4k-instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 17.45 GiB17.4 GB18.4 GB

+5.2%

(1.05x)

Pass

FM38

Phi-3-mini-4k-instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 10.91 GiB10.9 GB13.8 GB

+26.1%

(1.26x)

Miss

FM39

Phi-3-mini-4k-instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs2 seq1024: peak device memory 12.38 GiB12.4 GB13.3 GB

+7.4%

(1.07x)

Pass

FM40

Phi-3-mini-4k-instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 9.79 GiB9.8 GB12.7 GB

+29.2%

(1.29x)

Miss

FM41

Phi-3-mini-4k-instruct QLORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 5.13 GiB5.1 GB8.2 GB

+59.1%

(1.59x)

Miss

FM42

Phi-3-mini-4k-instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq4096: peak device memory 8.64 GiB8.6 GB9.0 GB

+4.3%

(1.04x)

Pass

FM43

Phi-3-mini-4k-instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 8.64 GiB8.6 GB8.9 GB

+3.6%

(1.04x)

Pass

FM44

Phi-3-mini-4k-instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 4.77 GiB4.8 GB5.5 GB

+16.4%

(1.16x)

Miss

FM45

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 50.11 GiB50.1 GB45.5 GB

-9.2%

(0.91x)

Pass

FM46

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 50.7 GiB50.7 GB50.9 GB

+0.4%

(1.00x)

Pass

FM47

Qwen2.5-14B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs4 seq2048: peak device memory 52.25 GiB52.3 GB52.2 GB

-0.2%

(1.00x)

Pass

FM48

Qwen2.5-14B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 34.37 GiB34.4 GB34.4 GB

+0.1%

(1.00x)

Pass

FM49

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq1024: peak device memory 41.59 GiB41.6 GB40.9 GB

-1.6%

(0.98x)

Pass

FM50

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 31.53 GiB31.5 GB34.5 GB

+9.4%

(1.09x)

Pass

FM51

Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 31.62 GiB31.6 GB36.8 GB

+16.4%

(1.16x)

Miss

FM52

Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 44.15 GiB44.1 GB47.5 GB

+7.6%

(1.08x)

Pass

FM53

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 32.39 GiB32.4 GB30.9 GB

-4.5%

(0.96x)

Pass

FM54

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 49.67 GiB49.7 GB44.4 GB

-10.6%

(0.89x)

Miss

FM55

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs4 seq1024: peak device memory 49.67 GiB49.7 GB44.4 GB

-10.6%

(0.89x)

Miss

FM56

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 53.9 GiB53.9 GB50.5 GB

-6.2%

(0.94x)

Pass

FM57

Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 37.88 GiB37.9 GB37.7 GB

-0.4%

(1.00x)

Pass

FM58

Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 59.43 GiB59.4 GB60.8 GB

+2.3%

(1.02x)

Pass

FM59

Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 36.38 GiB36.4 GB35.3 GB

-3.0%

(0.97x)

Pass

FM60

Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 49.55 GiB49.5 GB49.5 GB

-0.2%

(1.00x)

Pass

FM61

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 21.88 GiB21.9 GB22.1 GB

+0.8%

(1.01x)

Pass

FM62

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 39.11 GiB39.1 GB38.4 GB

-1.7%

(0.98x)

Pass

FM63

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 39.11 GiB39.1 GB37.4 GB

-4.4%

(0.96x)

Pass

FM64

Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 19.74 GiB19.7 GB19.0 GB

-3.6%

(0.96x)

Pass

FM65

Qwen2.5-14B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 29.98 GiB30.0 GB30.6 GB

+1.9%

(1.02x)

Pass

FM66

Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq8192: peak device memory 15.3 GiB15.3 GB14.9 GB

-2.7%

(0.97x)

Pass

FM67

Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq2048: peak device memory 13.02 GiB13.0 GB12.5 GB

-3.8%

(0.96x)

Pass

FM68

Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 21.08 GiB21.1 GB21.4 GB

+1.7%

(1.02x)

Pass

FM69

Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq4096: peak device memory 22.75 GiB22.8 GB23.1 GB

+1.5%

(1.02x)

Pass

FM70

Qwen2.5-3B-Instruct full FT, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 20.95 GiB20.9 GB33.7 GB

+61.0%

(1.61x)

Miss

FM71

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs1 seq16384: peak device memory 18.95 GiB18.9 GB19.5 GB

+3.1%

(1.03x)

Pass

FM72

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 16.3 GiB16.3 GB16.4 GB

+0.5%

(1.00x)

Pass

FM73

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 18.95 GiB18.9 GB18.0 GB

-4.9%

(0.95x)

Pass

FM74

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 18.95 GiB18.9 GB17.8 GB

-6.2%

(0.94x)

Pass

FM75

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 52.31 GiB52.3 GB45.5 GB

-13.0%

(0.87x)

Miss

FM76

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 51.45 GiB51.5 GB50.9 GB

-1.1%

(0.99x)

Pass

FM77

Qwen2.5-14B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs4 seq2048: peak device memory 52.06 GiB52.1 GB52.2 GB

+0.2%

(1.00x)

Pass

FM78

Qwen2.5-14B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs1 seq2048: peak device memory 34.64 GiB34.6 GB34.4 GB

-0.7%

(0.99x)

Pass

FM79

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs1 seq1024: peak device memory 41.81 GiB41.8 GB40.9 GB

-2.1%

(0.98x)

Pass

FM80

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 31.76 GiB31.8 GB34.5 GB

+8.6%

(1.09x)

Pass

FM81

Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 44.41 GiB44.4 GB47.5 GB

+7.0%

(1.07x)

Pass

FM82

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 49.93 GiB49.9 GB44.4 GB

-11.0%

(0.89x)

Miss

FM83

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 54.13 GiB54.1 GB50.5 GB

-6.6%

(0.93x)

Pass

FM84

Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 59.28 GiB59.3 GB60.8 GB

+2.6%

(1.03x)

Pass

FM85

Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 36.55 GiB36.5 GB35.3 GB

-3.5%

(0.97x)

Pass

FM86

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 22.06 GiB22.1 GB22.1 GB

+0.0%

(1.00x)

Pass

FM87

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 39.08 GiB39.1 GB37.4 GB

-4.4%

(0.96x)

Pass

FM88

Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 20.03 GiB20.0 GB19.0 GB

-5.0%

(0.95x)

Pass

FM89

Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq2048: peak device memory 13.33 GiB13.3 GB12.5 GB

-6.1%

(0.94x)

Pass

FM90

Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq4096: peak device memory 23.14 GiB23.1 GB23.1 GB

-0.2%

(1.00x)

Pass

FM91

Qwen2.5-3B-Instruct full FT, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 21.61 GiB21.6 GB33.7 GB

+56.0%

(1.56x)

Miss

FM92

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 16.6 GiB16.6 GB16.4 GB

-1.3%

(0.99x)

Pass

FM93

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 19.25 GiB19.3 GB18.0 GB

-6.4%

(0.94x)

Pass

FM94

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 19.25 GiB19.3 GB17.8 GB

-7.7%

(0.92x)

Pass

GM1

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 8192, 1x RTX 4090 (24GB): device memory 5.86 GiB (model 4403 + KV 1024 + compute 116 MiB)5.9 GB6.0 GB

+2.6%

(1.03x)

Pass

GM2

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 8.88 GiB (model 4403 + KV 4096 + compute 140 MiB)8.9 GB9.0 GB

+1.5%

(1.01x)

Pass

GM3

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 65536, 1x RTX 4090 (24GB): device memory 12.92 GiB (model 4403 + KV 8192 + compute 172 MiB)12.9 GB13.0 GB

+0.7%

(1.01x)

Pass

GM4

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 131072, 1x RTX 4090 (24GB): device memory 20.98 GiB (model 4403 + KV 16384 + compute 236 MiB)21.0 GB21.0 GB

+0.1%

(1.00x)

Pass

GM5

Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 9.27 GiB (model 8148 + KV 768 + compute 115 MiB)9.3 GB9.3 GB

+0.1%

(1.00x)

Pass

GM6

Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 16384, 1x RTX 4090 (24GB): device memory 11.53 GiB (model 8148 + KV 3072 + compute 127 MiB)11.5 GB11.5 GB

+0.0%

(1.00x)

Pass

GM7

Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 14.55 GiB (model 8148 + KV 6144 + compute 143 MiB)14.6 GB14.5 GB

-0.1%

(1.00x)

Pass

GM8

Qwen2.5-3B-Instruct Q8_0, llama_cpp 0.00.000.786 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 3.73 GiB (model 3128 + KV 144 + compute 81 MiB)3.7 GB3.7 GB

-1.1%

(0.99x)

Pass

GM9

Qwen2.5-3B-Instruct Q8_0, llama_cpp 0.00.000.786 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 4.74 GiB (model 3128 + KV 1152 + compute 109 MiB)4.7 GB4.7 GB

-1.5%

(0.99x)

Pass

GM10

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 4.88 GiB (model 4168 + KV 224 + compute 136 MiB)4.9 GB5.0 GB

+2.5%

(1.02x)

Pass

GM11

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 16384, 1x RTX 4090 (24GB): device memory 5.54 GiB (model 4168 + KV 896 + compute 148 MiB)5.5 GB5.7 GB

+2.0%

(1.02x)

Pass

GM12

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 6.44 GiB (model 4168 + KV 1792 + compute 164 MiB)6.4 GB6.5 GB

+1.4%

(1.01x)

Pass

GM13

Llama-3.1-8B-Instruct (llama3.1:8b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x RTX 4090 (24GB): device memory 9.02 GiB9.0 GB9.1 GB

+0.7%

(1.01x)

Pass

GM14

Llama-3.1-8B-Instruct (llama3.1:8b), ollama ollama version is 0.34.4 defaults, num_ctx 8192, 1x RTX 4090 (24GB): device memory 5.98 GiB6.0 GB6.1 GB

+1.7%

(1.02x)

Pass

GM15

Qwen2.5-14B-Instruct (qwen2.5:14b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 9.27 GiB9.3 GB9.4 GB

+1.4%

(1.01x)

Pass

GM16

Qwen2.5-14B-Instruct (qwen2.5:14b), ollama ollama version is 0.34.4 defaults, num_ctx 16384, 1x RTX 4090 (24GB): device memory 11.67 GiB11.7 GB11.7 GB

-0.2%

(1.00x)

Pass

GM17

Qwen2.5-3B-Instruct (qwen2.5:3b-instruct-q8_0), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 3.73 GiB3.7 GB3.7 GB

-0.5%

(0.99x)

Pass

GM18

Qwen2.5-3B-Instruct (qwen2.5:3b-instruct-q8_0), ollama ollama version is 0.34.4 defaults, num_ctx 32768, 1x RTX 4090 (24GB): device memory 4.85 GiB4.8 GB4.7 GB

-3.1%

(0.97x)

Pass

GM19

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x RTX 4090 (24GB): device memory 6.61 GiB6.6 GB6.6 GB

-0.3%

(1.00x)

Pass

GM20

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 4.88 GiB4.9 GB5.1 GB

+3.7%

(1.04x)

Pass

GM21

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 16384, 1x RTX 4090 (24GB): device memory 5.7 GiB5.7 GB5.7 GB

+0.2%

(1.00x)

Pass

GM22

Phi-3-mini-4k-instruct Q4_K_M, llama_cpp 0.00.001.232 I srv llama_server: initializing ... defaults, -c 2048, 1x RTX 4090 (24GB): device memory 3.44 GiB (model 2229 + KV 768 + compute 62 MiB)3.4 GB3.7 GB

+7.8%

(1.08x)

Pass

GM23

Phi-3-mini-4k-instruct Q4_K_M, llama_cpp 0.00.001.232 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 4.19 GiB (model 2229 + KV 1536 + compute 64 MiB)4.2 GB4.5 GB

+6.4%

(1.06x)

Pass

GM24

Phi-3-mini-4k-instruct Q8_0, llama_cpp 0.00.000.784 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 5.7 GiB (model 3773 + KV 1536 + compute 64 MiB)5.7 GB5.9 GB

+4.2%

(1.04x)

Pass

GM25

Phi-3-mini-4k-instruct (phi3:3.8b-mini-4k-instruct-q4_K_M), ollama ollama version is 0.34.4 defaults, default context (4096 for this VRAM tier), 1x RTX 4090 (24GB): device memory 4.19 GiB4.2 GB4.5 GB

+7.2%

(1.07x)

Pass

GM26

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.965 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 8.93 GiB (model 4403 + KV 4096 + compute 140 MiB)8.9 GB9.0 GB

+0.9%

(1.01x)

Pass

GM27

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.965 I srv llama_server: initializing ... defaults, -c 131072, 1x A100-SXM4 (80GB): device memory 21.02 GiB (model 4403 + KV 16384 + compute 236 MiB)21.0 GB21.0 GB

-0.0%

(1.00x)

Pass

GM28

Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.667 I srv llama_server: initializing ... defaults, -c 4096, 1x A100-SXM4 (80GB): device memory 19.77 GiB (model 18508 + KV 1024 + compute 196 MiB)19.8 GB19.3 GB

-2.5%

(0.97x)

Pass

GM29

Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.667 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 26.79 GiB (model 18508 + KV 8192 + compute 224 MiB)26.8 GB26.3 GB

-1.9%

(0.98x)

Pass

GM30

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.002.264 I srv llama_server: initializing ... defaults, -c 4096, 1x A100-SXM4 (80GB): device memory 4.91 GiB (model 4168 + KV 224 + compute 136 MiB)4.9 GB5.0 GB

+1.8%

(1.02x)

Pass

GM31

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.002.264 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 6.88 GiB (model 4168 + KV 1792 + compute 164 MiB)6.9 GB6.5 GB

-5.1%

(0.95x)

Pass

GM32

Qwen2.5-32B-Instruct (qwen2.5:32b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x A100-SXM4 (80GB): device memory 27.02 GiB27.0 GB26.5 GB

-1.8%

(0.98x)

Pass

GM33

Qwen2.5-32B-Instruct (qwen2.5:32b), ollama ollama version is 0.34.4 defaults, num_ctx 8192, 1x A100-SXM4 (80GB): device memory 20.97 GiB21.0 GB20.5 GB

-2.1%

(0.98x)

Pass

GM34

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x A100-SXM4 (80GB): device memory 6.64 GiB6.6 GB6.6 GB

-0.8%

(0.99x)

Pass

GM35

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x A100-SXM4 (80GB): device memory 4.89 GiB4.9 GB5.1 GB

+3.5%

(1.03x)

Pass

GM36

Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.557 I srv llama_server: initializing ... defaults, -c 8192, 1x H100 SXM (80GB): device memory 20.88 GiB (model 18508 + KV 2048 + compute 200 MiB)20.9 GB20.3 GB

-2.9%

(0.97x)

Pass

GM37

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.683 I srv llama_server: initializing ... defaults, -c 4096, 1x H100 SXM (80GB): device memory 5.03 GiB (model 4168 + KV 224 + compute 136 MiB)5.0 GB5.0 GB

-0.6%

(0.99x)

Pass

GM38

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.683 I srv llama_server: initializing ... defaults, -c 32768, 1x H100 SXM (80GB): device memory 6.58 GiB (model 4168 + KV 1792 + compute 164 MiB)6.6 GB6.5 GB

-0.8%

(0.99x)

Pass

GM39

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x H100 SXM (80GB): device memory 5.03 GiB5.0 GB5.1 GB

+0.6%

(1.01x)

Pass

GM40

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 32768, 1x H100 SXM (80GB): device memory 6.75 GiB6.8 GB6.6 GB

-2.4%

(0.98x)

Pass

TM1

Qwen2.5-3B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (12733 x 32 tokens, free_gpu_memory_fraction 0.9)12733.0 GB12296.0 GB

-3.4%

(0.97x)

Pass

TM2

Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (3458 x 32 tokens, free_gpu_memory_fraction 0.9)3458.0 GB3465.0 GB

+0.2%

(1.00x)

Pass

TM3

Phi-3-mini-4k-instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (1116 x 32 tokens, free_gpu_memory_fraction 0.9)1116.0 GB1056.0 GB

-5.4%

(0.95x)

Pass

TM4

Qwen2.5-14B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (7561 x 32 tokens, free_gpu_memory_fraction 0.9)7561.0 GB7550.0 GB

-0.1%

(1.00x)

Pass

TM5

Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (32853 x 32 tokens, free_gpu_memory_fraction 0.9)32853.0 GB32872.0 GB

+0.1%

(1.00x)

Pass

TM6

Llama-3.1-8B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (14132 x 32 tokens, free_gpu_memory_fraction 0.9)14132.0 GB14192.0 GB

+0.4%

(1.00x)

Pass

TM7

Qwen2.5-14B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (7475 x 32 tokens, free_gpu_memory_fraction 0.9)7475.0 GB7550.0 GB

+1.0%

(1.01x)

Pass

TM8

Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (32498 x 32 tokens, free_gpu_memory_fraction 0.9)32498.0 GB32872.0 GB

+1.2%

(1.01x)

Pass

VM9

Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (65412 x 16 tokens, gpu_memory_utilization 0.92)65412.0 GB66519.0 GB

+1.7%

(1.02x)

Pass