ApX 标志ApX 标志

趋近智

筛选

架构类型

数据来源

状态

精度评分卡

对比预测值与已验证的真实基准配置。

打开计算器

校准总览

测试套件带内通过率中位数误差GMFE 因子误差系统偏差最差情况
显存占用92.6%1.8%1.05x

无偏差

E46 (1.90x)
解码速度88.1%8.4%1.13x

微偏低

T9 (0.30x)
微调速度90.4%6.2%1.11x

微偏低

T22 (0.55x)

内存预测具有高度确定性且非常准确,而速度估算对运行和硬件配置更为敏感。未命中主要是由于源基准中未指定参数数据所致。

Showing 270 of 270 cases

ID场景与配置测量值预测值比例状态

A1

Llama-3-8B (8.03B) fp16 weights = safetensors size15.0 GB15.0 GB

-0.1%

(1.00x)

通过

A2

Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size18.5 GB18.2 GB

-1.7%

(0.98x)

通过

A3

Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size25.1 GB25.3 GB

+1.1%

(1.01x)

通过

A4

Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size32.4 GB31.0 GB

-4.4%

(0.96x)

通过

A5

Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size44.1 GB42.2 GB

-4.3%

(0.96x)

通过

A6

Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once24.6 GB25.7 GB

+4.2%

(1.04x)

通过

A7

Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size13.5 GB13.5 GB

+0.0%

(1.00x)

通过

A8

Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab)6.0 GB6.0 GB

-0.2%

(1.00x)

通过

A9

Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit)4.4 GB4.4 GB

+1.8%

(1.02x)

通过

A10

Llama-2-7B (6.74B) fp16 weights = safetensors size12.6 GB12.6 GB

+0.0%

(1.00x)

通过

B1

Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128)1.0 GB1.0 GB

+0.0%

(1.00x)

通过

B2

Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128)4.0 GB4.0 GB

+0.0%

(1.00x)

通过

B3

Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128)1.3 GB1.3 GB

+0.0%

(1.00x)

通过

B4

DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer)2.1 GB2.1 GB

+0.0%

(1.00x)

通过

B4b

Kimi-K3 hybrid MLA + Linear Attention KV cache, 1M ctx, fp16 (93L, 576 cached dims/MLA layer, 74.2% linear reduction)27.0 GB27.0 GB

+0.0%

(1.00x)

通过

B5

Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128)4.0 GB4.0 GB

+0.0%

(1.00x)

通过

B6

Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128)2.5 GB2.5 GB

+0.0%

(1.00x)

通过

B7

Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128)2.0 GB2.0 GB

+0.0%

(1.00x)

通过

B8

Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128)1.8 GB1.8 GB

+0.0%

(1.00x)

通过

B9

Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress16.0 GB16.0 GB

+0.0%

(1.00x)

通过

D1

65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB)44.0 GB42.9 GB

-2.6%

(0.97x)

通过

D2

Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth)8.5 GB7.4 GB

-12.8%

(0.87x)

未通过

D3

Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB)48.0 GB42.7 GB

-11.0%

(0.89x)

未通过

D4

7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations)104.3 GB122.3 GB

+17.3%

(1.17x)

通过

D5

70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading70.0 GB47.0 GB

-32.8%

(0.67x)

通过

D7

Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100)21.3 GB21.2 GB

-0.5%

(1.00x)

通过

E1

End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded)17.5 GB16.3 GB

-6.6%

(0.93x)

未通过

E2

End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB)24.0 GB23.2 GB

-3.3%

(0.97x)

通过

A11

Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight377.2 GB378.3 GB

+0.3%

(1.00x)

通过

A12

Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight754.4 GB754.4 GB

+0.0%

(1.00x)

通过

B10

Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress31.5 GB31.5 GB

+0.0%

(1.00x)

通过

B11

DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims)9.6 GB9.6 GB

+0.0%

(1.00x)

通过

B12

DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims)4.3 GB4.3 GB

+0.0%

(1.00x)

通过

B18

DeepSeek-V4.1-Flash KV cache, 1M ctx, fp16 (40L CED, CSA2 cross-layer reuse)3.4 GB3.4 GB

+0.0%

(1.00x)

通过

B19

DeepSeek-V4.1-Flash KV cache, 1M ctx, nvfp4 (40L CED, CSA2 890 B/tok)0.9 GB0.9 GB

+0.0%

(1.00x)

通过

D8

Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload)54.7 GB45.7 GB

-16.4%

(0.84x)

通过

E3

End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM)486.8 GB417.2 GB

-14.3%

(0.86x)

未通过

E4

End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM)881.2 GB800.3 GB

-9.2%

(0.91x)

通过

D9

Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100)14.2 GB13.8 GB

-2.7%

(0.97x)

通过

D10

Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU)36.7 GB31.0 GB

-15.4%

(0.85x)

通过

D11

Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB)12.5 GB13.1 GB

+4.4%

(1.04x)

通过

D12

Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling21.9 GB16.9 GB

-23.1%

(0.77x)

通过

D13

Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB)17.4 GB20.9 GB

+20.3%

(1.20x)

通过

D14

Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB)7.9 GB10.3 GB

+30.0%

(1.30x)

通过

D15

Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU)29.6 GB24.9 GB

-15.7%

(0.84x)

通过

D16

Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU)44.8 GB45.3 GB

+1.1%

(1.01x)

通过

D17

Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB)24.9 GB22.9 GB

-7.9%

(0.92x)

通过

D18

33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B21.0 GB21.7 GB

+3.3%

(1.03x)

通过

E5

PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx8.3 GB6.6 GB

-20.8%

(0.79x)

未通过

E6

PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx31.6 GB26.4 GB

-16.5%

(0.84x)

未通过

E7

PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx475.2 GB319.4 GB

-32.8%

(0.67x)

未通过

E9

PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx215.4 GB209.8 GB

-2.6%

(0.97x)

通过

A13

Qwen3-Omni-30B-A3B (33.0B total) BF16 weights = safetensors size61.5 GB61.5 GB

+0.0%

(1.00x)

通过

A14

Qwen3-Omni-30B-A3B (33.0B total, 3B aux) W4A16 weights = AutoRound int4 size23.3 GB22.2 GB

-4.8%

(0.95x)

通过

E10

End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE)2100.0 GB2073.6 GB

-1.3%

(0.99x)

通过

B13

Llama-3.1-70B KV Cache with batch size 3280.0 GB80.0 GB

+0.0%

(1.00x)

通过

D19

DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes)76.5 GB80.3 GB

+4.9%

(1.05x)

通过

A15

Inkling-Small (276.0B / 12.0B active) BF16 weights = base safetensors size514.1 GB514.1 GB

+0.0%

(1.00x)

通过

A16

Inkling-Small (276.0B / 12.0B active) NVFP4 weights130.4 GB130.8 GB

+0.3%

(1.00x)

通过

A17

Nemotron 3.5 Lightning (30.0B / 3.0B active) BF16 weights = safetensors size55.9 GB55.9 GB

+0.0%

(1.00x)

通过

A18

Nemotron 3.5 Lightning (30.0B / 3.0B active) NVFP4 weights14.9 GB14.9 GB

+0.3%

(1.00x)

通过

A19

Qwen3.8-2.4T-A95B (2400.0B / 95.0B active) NVFP4 weights1117.6 GB1123.3 GB

+0.5%

(1.01x)

通过

B14

Inkling-Small KV cache, 8k ctx, fp16 (42L, 8 KV heads, head 128)1.3 GB1.3 GB

+0.0%

(1.00x)

通过

B15

Nemotron 3.5 Lightning KV cache, 8k ctx, fp16 (52L, hidden 2688, linear attention ratio 0.75)1.1 GB1.1 GB

+0.0%

(1.00x)

通过

A20

Muse-Glimmer-30B (29.6B) bf16 weights = safetensors size55.1 GB55.1 GB

+0.0%

(1.00x)

通过

B16

Muse-Glimmer-30B KV cache, 8k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048)0.2 GB0.2 GB

+0.0%

(1.00x)

通过

B17

Muse-Glimmer-30B KV cache, 128k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048)1.7 GB1.7 GB

+0.0%

(1.00x)

通过

E12

End-to-end: Muse-Glimmer-30B Q4_K_M inference, 8k ctx (claimed ~17-24GB on RTX 4090/5090)19.5 GB20.8 GB

+6.6%

(1.07x)

通过

E13

Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 1 (vLLM memory profiling: ~140GB weights + ~42GB KV + ~19GB framework/NCCL)203.0 GB205.5 GB

+1.2%

(1.01x)

通过

E14

Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 4 (vLLM memory profiling with chunked prefill: ~205GB total across cluster)205.5 GB205.5 GB

-0.0%

(1.00x)

通过

A21

Llama-3.2-11B-Vision (10.67B total, 2.64B aux) BF16 weights = safetensors size19.9 GB19.9 GB

+0.0%

(1.00x)

通过

A22

Qwen2-VL-7B-Instruct (8.29B untied safetensors) BF16 weights = safetensors size15.4 GB15.4 GB

+0.0%

(1.00x)

通过

A23

Qwen2-VL-7B-Instruct-GPTQ-Int4 weights = GPTQ Int4 safetensors size6.5 GB6.3 GB

-1.9%

(0.98x)

通过

A24

Qwen2.5-3B-Instruct (3.09B) BF16 weights = published safetensors size5.8 GB5.8 GB

+0.2%

(1.00x)

通过

A25

Mistral-7B-Instruct-v0.3 (7.25B) BF16 weights = published safetensors size13.5 GB13.5 GB

+0.0%

(1.00x)

通过

A26

Qwen2.5-0.5B-Instruct (0.49B) BF16 weights = published single safetensors size0.9 GB0.9 GB

-1.1%

(0.99x)

通过

A27

Qwen3.5-0.8B (0.87B total with 0.10B visual encoder) BF16 weights = published safetensors size1.6 GB1.6 GB

-0.6%

(0.99x)

通过

A28

Gemma 4 E2B (5.12B total with 0.49B vision+audio) BF16 weights = published safetensors size9.6 GB9.5 GB

-0.1%

(1.00x)

通过

B20

Qwen2.5-3B KV cache, 8k ctx, fp16 (36L, 2 KV heads, head 128)0.3 GB0.3 GB

+0.0%

(1.00x)

通过

E15

Qwen2.5-3B-Instruct FP16 raw forward (all-position logits), RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 6.85GB)6.8 GB7.3 GB

+7.0%

(1.07x)

通过

E16

Qwen2.5-3B-Instruct FP16 raw forward (all-position logits), RTX 4090 (24GB), 8k ctx, batch 4 (measured peak: 19.773GB; fits - the earlier OOM was the harness holding two logits tensors)19.8 GB19.7 GB

-0.6%

(0.99x)

通过

E17

DeepSeek-V2-Lite (16B MoE total / 2.4B active) FP16 weights loading on RTX 4090 (24GB) -> OOMInsufficientInsufficient-
通过

E18

Qwen2-VL-2B (2.21B total with auxiliary visual encoder) BF16 empirical weights on RTX 40904.1 GB4.1 GB

+0.0%

(1.00x)

通过

E19

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 15.41GB)15.4 GB15.9 GB

+3.2%

(1.03x)

通过

E20

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 4 (measured peak: 18.59GB)18.6 GB18.7 GB

+0.4%

(1.00x)

通过

E21

Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 8 (measured peak: 23.05GB)23.1 GB22.8 GB

-1.0%

(0.99x)

通过

E22

Qwen2.5-7B-Instruct FP16 empirical model weights on RTX 409014.2 GB14.2 GB

+0.1%

(1.00x)

通过

E23

Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 1k ctx, batch 1 (measured peak: 18.40GB)18.4 GB19.3 GB

+5.0%

(1.05x)

通过

E24

Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 19.87GB)19.9 GB19.7 GB

-0.9%

(0.99x)

通过

E25

Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 1 (measured peak: 17.19GB)17.2 GB17.2 GB

+0.2%

(1.00x)

通过

E26

Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 8 (measured peak: 32.73GB)32.7 GB30.1 GB

-8.1%

(0.92x)

通过

E27

Llama-3.1-8B-Instruct QLoRA r=16 training, 1x RTX 4090 (24GB), bs=2 seq=2048 (measured peak: 19.71GB)19.7 GB19.7 GB

-0.2%

(1.00x)

通过

E28

Qwen2.5-3B-Instruct full FT, 1x A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 33.59GB)33.6 GB32.1 GB

-4.3%

(0.96x)

通过

E29

DeepSeek-R1-Distill-Qwen-14B FP16 empirical model weights, A100-SXM4 (80GB)27.5 GB27.5 GB

+0.0%

(1.00x)

通过

E30

DeepSeek-R1-Distill-Qwen-14B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 30.30GB)30.3 GB30.5 GB

+0.7%

(1.01x)

通过

E31

DeepSeek-R1-Distill-Qwen-7B FP16 empirical model weights, RTX 4090 (24GB)14.2 GB14.2 GB

+0.1%

(1.00x)

通过

E32

DeepSeek-R1-Distill-Qwen-7B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 16.23GB)16.2 GB15.9 GB

-2.0%

(0.98x)

通过

E33

DeepSeek-R1-Distill-Qwen-32B FP16 empirical model weights, A100-SXM4 (80GB)61.0 GB61.0 GB

-0.0%

(1.00x)

通过

E34

DeepSeek-R1-Distill-Qwen-32B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 66.65GB)66.7 GB66.7 GB

+0.0%

(1.00x)

通过

E35

Qwen2.5-3B-Instruct FP16 inference, H200 (141GB), 2k ctx, batch 1 (measured peak: 6.76GB)6.8 GB6.8 GB

+0.0%

(1.00x)

通过

E36

Qwen3.8-27B FP16 empirical model weights, A100-SXM4 (80GB)50.1 GB50.1 GB

+0.0%

(1.00x)

通过

E37

Qwen3.8-27B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 54.67GB)54.7 GB54.7 GB

-0.0%

(1.00x)

通过

E38

Ministral-3-8B-Instruct-2512 FP16 empirical model weights, A100-SXM4 (80GB)16.6 GB16.6 GB

+0.0%

(1.00x)

通过

E39

Ministral-3-8B-Instruct-2512 FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 20.38GB)20.4 GB18.7 GB

-8.4%

(0.92x)

通过

E40

Granite-4.2-8B FP16 empirical model weights, RTX 4090 (24GB)16.4 GB16.4 GB

+0.0%

(1.00x)

通过

E41

Granite-4.2-8B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 18.89GB)18.9 GB18.5 GB

-2.3%

(0.98x)

通过

E42

Gemma-4-E2B FP16 empirical model weights, RTX 4090 (24GB)9.5 GB9.5 GB

-0.1%

(1.00x)

通过

E43

Gemma-4-E2B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 12.81GB)12.8 GB12.4 GB

-2.8%

(0.97x)

通过

E44

Gemma-4-12B FP16 empirical model weights, L40S (48GB)22.3 GB22.3 GB

-0.1%

(1.00x)

通过

E45

Gemma-4-12B FP16 inference, L40S (48GB), 2k ctx, batch 1 (measured peak: 26.62GB)26.6 GB25.9 GB

-2.8%

(0.97x)

通过

E46

Ministral-3-3B Full FT 32-bit AdamW, A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 36.95GB)37.0 GB70.1 GB

+89.7%

(1.90x)

未通过

E47

DeepSeek-R1-Distill-Qwen-7B LoRA 32-bit AdamW training, RTX 4090 (24GB) -> OOMInsufficientInsufficient-
通过

E48

Ministral-3-8B-Instruct-2512 LoRA 8-bit AdamW training, RTX 4090 (24GB) -> OOMInsufficientInsufficient-
通过

E49

Qwen3.8-Flash-Next (176.9B) FP16 inference, 2x H200 (141GB) -> OOM (weights exceed cluster capacity)InsufficientInsufficient-
通过

E50

Qwen2.5-3B-Instruct FP16 prefill forward with last-token logits (as generate() does), RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 6.27GB)6.3 GB6.8 GB

+7.8%

(1.08x)

通过

E51

Qwen2.5-3B-Instruct FP16 prefill forward with last-token logits (as generate() does), RTX 4090 (24GB), 8k ctx, batch 4 (measured peak: 10.5GB)10.5 GB10.4 GB

-1.0%

(0.99x)

通过

E52

Gemma-4-12B BF16 inference, 2x B300 SXM6 (288GB), 2k ctx, batch 1 (measured physical peak: 28.72GB)28.7 GB26.3 GB

-8.5%

(0.92x)

通过

E53

Devstral-2-123B-Instruct FP8 inference, 2x B300 SXM6 (288GB), 2k ctx, batch 1 (measured physical peak: 128.15GB)128.2 GB139.8 GB

+9.1%

(1.09x)

通过

E54

Qwen2.5-7B-Instruct FP16 eager inference, AMD Instinct MI350X (288GB), 2k ctx, batch 1 (measured physical peak: 16.03GB)16.0 GB15.9 GB

-0.9%

(0.99x)

通过

VM1

Qwen2.5-32B-Instruct BF16, vLLM 0.30.0 defaults, 2x H100 SXM (80GB) TP=2: KV cache blocks after profiling (19691 x 16 tokens, gpu_memory_utilization 0.92)19691.0 GB19847.0 GB

+0.8%

(1.01x)

通过

VM2

Qwen2.5-3B-Instruct BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (25334 x 16 tokens, gpu_memory_utilization 0.92)25334.0 GB25158.0 GB

-0.7%

(0.99x)

通过

VM3

Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (5422 x 16 tokens, gpu_memory_utilization 0.92)5422.0 GB6225.0 GB

+14.8%

(1.15x)

通过

VM4

DeepSeek-R1-Distill-Qwen-7B BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (6567 x 16 tokens, gpu_memory_utilization 0.92)6567.0 GB6225.0 GB

-5.2%

(0.95x)

通过

VM5

Qwen2.5-14B-Instruct BF16, vLLM 0.30.0 defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (14195 x 16 tokens, gpu_memory_utilization 0.92)14195.0 GB14824.0 GB

+4.4%

(1.04x)

通过

VM6

Qwen2.5-14B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (13630 x 16 tokens, gpu_memory_utilization 0.92)13630.0 GB14704.0 GB

+7.9%

(1.08x)

通过

VM7

Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (64060 x 16 tokens, gpu_memory_utilization 0.92)64060.0 GB66226.0 GB

+3.4%

(1.03x)

通过

VM8

Llama-3.1-8B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (27701 x 16 tokens, gpu_memory_utilization 0.92)27701.0 GB28554.0 GB

+3.1%

(1.03x)

通过

FM1

Llama-3.1-8B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 14.07 GiB14.1 GB15.4 GB

+9.2%

(1.09x)

通过

FM2

Llama-3.1-8B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOMInsufficientInsufficient-
通过

FM3

Llama-3.1-8B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 15.06 GiB15.1 GB16.9 GB

+11.9%

(1.12x)

未通过

FM4

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 13.87 GiB13.9 GB14.2 GB

+2.5%

(1.02x)

通过

FM5

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 21.11 GiB21.1 GB19.6 GB

-7.3%

(0.93x)

通过

FM6

Qwen2.5-3B-Instruct LORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOMInsufficientInsufficient-
通过

FM7

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 16.45 GiB16.4 GB16.4 GB

-0.4%

(1.00x)

通过

FM8

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 22.82 GiB22.8 GB23.9 GB

+4.8%

(1.05x)

通过

FM9

Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs4 seq512: peak device memory 21.11 GiB21.1 GB19.6 GB

-7.3%

(0.93x)

通过

FM10

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 19.91 GiB19.9 GB21.6 GB

+8.5%

(1.08x)

通过

FM11

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 14.43 GiB14.4 GB15.1 GB

+4.8%

(1.05x)

通过

FM12

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 22.69 GiB22.7 GB21.9 GB

-3.7%

(0.96x)

通过

FM13

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOMInsufficientInsufficient-
通过

FM14

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 15.7 GiB15.7 GB17.1 GB

+9.2%

(1.09x)

通过

FM15

Qwen2.5-7B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs4 seq2048 -> OOMInsufficientInsufficient-
通过

FM16

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 10.32 GiB10.3 GB10.7 GB

+3.5%

(1.03x)

通过

FM17

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 12.88 GiB12.9 GB12.9 GB

-0.1%

(1.00x)

通过

FM18

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 18.01 GiB18.0 GB17.8 GB

-1.3%

(0.99x)

通过

FM19

Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 18.01 GiB18.0 GB17.3 GB

-4.2%

(0.96x)

通过

FM20

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 18.87 GiB18.9 GB19.5 GB

+3.3%

(1.03x)

通过

FM21

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 21.83 GiB21.8 GB22.1 GB

+1.1%

(1.01x)

通过

FM22

Qwen2.5-7B-Instruct LORA r=16, torchtune None defaults, 1x RTX 4090 (24GB), bs2 seq4096 -> OOMInsufficientInsufficient-
通过

FM23

Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 16.46 GiB16.5 GB17.0 GB

+3.2%

(1.03x)

通过

FM24

Llama-3.1-8B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 8.52 GiB8.5 GB8.6 GB

+0.6%

(1.01x)

通过

FM25

Llama-3.1-8B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.55 GiB7.5 GB7.7 GB

+1.7%

(1.02x)

通过

FM26

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 7.08 GiB7.1 GB7.3 GB

+3.7%

(1.04x)

通过

FM27

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 7.94 GiB7.9 GB8.3 GB

+3.9%

(1.04x)

通过

FM28

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.55 GiB7.5 GB7.5 GB

-0.4%

(1.00x)

通过

FM29

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 7.94 GiB7.9 GB8.0 GB

+0.8%

(1.01x)

通过

FM30

Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 7.94 GiB7.9 GB7.9 GB

-0.9%

(0.99x)

通过

FM31

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 16.26 GiB16.3 GB16.4 GB

+0.7%

(1.01x)

通过

FM32

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 6.89 GiB6.9 GB7.2 GB

+3.9%

(1.04x)

通过

FM33

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 8.01 GiB8.0 GB8.2 GB

+2.7%

(1.03x)

通过

FM34

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.38 GiB7.4 GB7.4 GB

+0.1%

(1.00x)

通过

FM35

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 8.01 GiB8.0 GB8.0 GB

-0.4%

(1.00x)

通过

FM36

Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 8.01 GiB8.0 GB7.8 GB

-2.0%

(0.98x)

通过

FM37

Phi-3-mini-4k-instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 17.45 GiB17.4 GB18.4 GB

+5.2%

(1.05x)

通过

FM38

Phi-3-mini-4k-instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 10.91 GiB10.9 GB13.8 GB

+26.1%

(1.26x)

未通过

FM39

Phi-3-mini-4k-instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs2 seq1024: peak device memory 12.38 GiB12.4 GB13.3 GB

+7.4%

(1.07x)

通过

FM40

Phi-3-mini-4k-instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 9.79 GiB9.8 GB12.7 GB

+29.2%

(1.29x)

未通过

FM41

Phi-3-mini-4k-instruct QLORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 5.13 GiB5.1 GB8.2 GB

+59.1%

(1.59x)

未通过

FM42

Phi-3-mini-4k-instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq4096: peak device memory 8.64 GiB8.6 GB9.0 GB

+4.3%

(1.04x)

通过

FM43

Phi-3-mini-4k-instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 8.64 GiB8.6 GB8.9 GB

+3.6%

(1.04x)

通过

FM44

Phi-3-mini-4k-instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 4.77 GiB4.8 GB5.5 GB

+16.4%

(1.16x)

未通过

FM45

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 50.11 GiB50.1 GB45.5 GB

-9.2%

(0.91x)

通过

FM46

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 50.7 GiB50.7 GB50.9 GB

+0.4%

(1.00x)

通过

FM47

Qwen2.5-14B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs4 seq2048: peak device memory 52.25 GiB52.3 GB52.2 GB

-0.2%

(1.00x)

通过

FM48

Qwen2.5-14B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 34.37 GiB34.4 GB34.4 GB

+0.1%

(1.00x)

通过

FM49

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq1024: peak device memory 41.59 GiB41.6 GB40.9 GB

-1.6%

(0.98x)

通过

FM50

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 31.53 GiB31.5 GB34.5 GB

+9.4%

(1.09x)

通过

FM51

Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 31.62 GiB31.6 GB36.8 GB

+16.4%

(1.16x)

未通过

FM52

Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 44.15 GiB44.1 GB47.5 GB

+7.6%

(1.08x)

通过

FM53

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 32.39 GiB32.4 GB30.9 GB

-4.5%

(0.96x)

通过

FM54

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 49.67 GiB49.7 GB44.4 GB

-10.6%

(0.89x)

未通过

FM55

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs4 seq1024: peak device memory 49.67 GiB49.7 GB44.4 GB

-10.6%

(0.89x)

未通过

FM56

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 53.9 GiB53.9 GB50.5 GB

-6.2%

(0.94x)

通过

FM57

Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 37.88 GiB37.9 GB37.7 GB

-0.4%

(1.00x)

通过

FM58

Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 59.43 GiB59.4 GB60.8 GB

+2.3%

(1.02x)

通过

FM59

Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 36.38 GiB36.4 GB35.3 GB

-3.0%

(0.97x)

通过

FM60

Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 49.55 GiB49.5 GB49.5 GB

-0.2%

(1.00x)

通过

FM61

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 21.88 GiB21.9 GB22.1 GB

+0.8%

(1.01x)

通过

FM62

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 39.11 GiB39.1 GB38.4 GB

-1.7%

(0.98x)

通过

FM63

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 39.11 GiB39.1 GB37.4 GB

-4.4%

(0.96x)

通过

FM64

Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 19.74 GiB19.7 GB19.0 GB

-3.6%

(0.96x)

通过

FM65

Qwen2.5-14B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 29.98 GiB30.0 GB30.6 GB

+1.9%

(1.02x)

通过

FM66

Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq8192: peak device memory 15.3 GiB15.3 GB14.9 GB

-2.7%

(0.97x)

通过

FM67

Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq2048: peak device memory 13.02 GiB13.0 GB12.5 GB

-3.8%

(0.96x)

通过

FM68

Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 21.08 GiB21.1 GB21.4 GB

+1.7%

(1.02x)

通过

FM69

Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq4096: peak device memory 22.75 GiB22.8 GB23.1 GB

+1.5%

(1.02x)

通过

FM70

Qwen2.5-3B-Instruct full FT, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 20.95 GiB20.9 GB33.7 GB

+61.0%

(1.61x)

未通过

FM71

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs1 seq16384: peak device memory 18.95 GiB18.9 GB19.5 GB

+3.1%

(1.03x)

通过

FM72

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 16.3 GiB16.3 GB16.4 GB

+0.5%

(1.00x)

通过

FM73

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 18.95 GiB18.9 GB18.0 GB

-4.9%

(0.95x)

通过

FM74

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 18.95 GiB18.9 GB17.8 GB

-6.2%

(0.94x)

通过

FM75

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 52.31 GiB52.3 GB45.5 GB

-13.0%

(0.87x)

未通过

FM76

Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 51.45 GiB51.5 GB50.9 GB

-1.1%

(0.99x)

通过

FM77

Qwen2.5-14B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs4 seq2048: peak device memory 52.06 GiB52.1 GB52.2 GB

+0.2%

(1.00x)

通过

FM78

Qwen2.5-14B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs1 seq2048: peak device memory 34.64 GiB34.6 GB34.4 GB

-0.7%

(0.99x)

通过

FM79

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs1 seq1024: peak device memory 41.81 GiB41.8 GB40.9 GB

-2.1%

(0.98x)

通过

FM80

Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 31.76 GiB31.8 GB34.5 GB

+8.6%

(1.09x)

通过

FM81

Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 44.41 GiB44.4 GB47.5 GB

+7.0%

(1.07x)

通过

FM82

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 49.93 GiB49.9 GB44.4 GB

-11.0%

(0.89x)

未通过

FM83

Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 54.13 GiB54.1 GB50.5 GB

-6.6%

(0.93x)

通过

FM84

Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 59.28 GiB59.3 GB60.8 GB

+2.6%

(1.03x)

通过

FM85

Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 36.55 GiB36.5 GB35.3 GB

-3.5%

(0.97x)

通过

FM86

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 22.06 GiB22.1 GB22.1 GB

+0.0%

(1.00x)

通过

FM87

Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 39.08 GiB39.1 GB37.4 GB

-4.4%

(0.96x)

通过

FM88

Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 20.03 GiB20.0 GB19.0 GB

-5.0%

(0.95x)

通过

FM89

Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq2048: peak device memory 13.33 GiB13.3 GB12.5 GB

-6.1%

(0.94x)

通过

FM90

Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq4096: peak device memory 23.14 GiB23.1 GB23.1 GB

-0.2%

(1.00x)

通过

FM91

Qwen2.5-3B-Instruct full FT, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 21.61 GiB21.6 GB33.7 GB

+56.0%

(1.56x)

未通过

FM92

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 16.6 GiB16.6 GB16.4 GB

-1.3%

(0.99x)

通过

FM93

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 19.25 GiB19.3 GB18.0 GB

-6.4%

(0.94x)

通过

FM94

Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 19.25 GiB19.3 GB17.8 GB

-7.7%

(0.92x)

通过

GM1

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 8192, 1x RTX 4090 (24GB): device memory 5.86 GiB (model 4403 + KV 1024 + compute 116 MiB)5.9 GB6.0 GB

+2.6%

(1.03x)

通过

GM2

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 8.88 GiB (model 4403 + KV 4096 + compute 140 MiB)8.9 GB9.0 GB

+1.5%

(1.01x)

通过

GM3

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 65536, 1x RTX 4090 (24GB): device memory 12.92 GiB (model 4403 + KV 8192 + compute 172 MiB)12.9 GB13.0 GB

+0.7%

(1.01x)

通过

GM4

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 131072, 1x RTX 4090 (24GB): device memory 20.98 GiB (model 4403 + KV 16384 + compute 236 MiB)21.0 GB21.0 GB

+0.1%

(1.00x)

通过

GM5

Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 9.27 GiB (model 8148 + KV 768 + compute 115 MiB)9.3 GB9.3 GB

+0.1%

(1.00x)

通过

GM6

Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 16384, 1x RTX 4090 (24GB): device memory 11.53 GiB (model 8148 + KV 3072 + compute 127 MiB)11.5 GB11.5 GB

+0.0%

(1.00x)

通过

GM7

Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 14.55 GiB (model 8148 + KV 6144 + compute 143 MiB)14.6 GB14.5 GB

-0.1%

(1.00x)

通过

GM8

Qwen2.5-3B-Instruct Q8_0, llama_cpp 0.00.000.786 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 3.73 GiB (model 3128 + KV 144 + compute 81 MiB)3.7 GB3.7 GB

-1.1%

(0.99x)

通过

GM9

Qwen2.5-3B-Instruct Q8_0, llama_cpp 0.00.000.786 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 4.74 GiB (model 3128 + KV 1152 + compute 109 MiB)4.7 GB4.7 GB

-1.5%

(0.99x)

通过

GM10

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 4.88 GiB (model 4168 + KV 224 + compute 136 MiB)4.9 GB5.0 GB

+2.5%

(1.02x)

通过

GM11

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 16384, 1x RTX 4090 (24GB): device memory 5.54 GiB (model 4168 + KV 896 + compute 148 MiB)5.5 GB5.7 GB

+2.0%

(1.02x)

通过

GM12

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 6.44 GiB (model 4168 + KV 1792 + compute 164 MiB)6.4 GB6.5 GB

+1.4%

(1.01x)

通过

GM13

Llama-3.1-8B-Instruct (llama3.1:8b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x RTX 4090 (24GB): device memory 9.02 GiB9.0 GB9.1 GB

+0.7%

(1.01x)

通过

GM14

Llama-3.1-8B-Instruct (llama3.1:8b), ollama ollama version is 0.34.4 defaults, num_ctx 8192, 1x RTX 4090 (24GB): device memory 5.98 GiB6.0 GB6.1 GB

+1.7%

(1.02x)

通过

GM15

Qwen2.5-14B-Instruct (qwen2.5:14b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 9.27 GiB9.3 GB9.4 GB

+1.4%

(1.01x)

通过

GM16

Qwen2.5-14B-Instruct (qwen2.5:14b), ollama ollama version is 0.34.4 defaults, num_ctx 16384, 1x RTX 4090 (24GB): device memory 11.67 GiB11.7 GB11.7 GB

-0.2%

(1.00x)

通过

GM17

Qwen2.5-3B-Instruct (qwen2.5:3b-instruct-q8_0), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 3.73 GiB3.7 GB3.7 GB

-0.5%

(0.99x)

通过

GM18

Qwen2.5-3B-Instruct (qwen2.5:3b-instruct-q8_0), ollama ollama version is 0.34.4 defaults, num_ctx 32768, 1x RTX 4090 (24GB): device memory 4.85 GiB4.8 GB4.7 GB

-3.1%

(0.97x)

通过

GM19

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x RTX 4090 (24GB): device memory 6.61 GiB6.6 GB6.6 GB

-0.3%

(1.00x)

通过

GM20

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 4.88 GiB4.9 GB5.1 GB

+3.7%

(1.04x)

通过

GM21

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 16384, 1x RTX 4090 (24GB): device memory 5.7 GiB5.7 GB5.7 GB

+0.2%

(1.00x)

通过

GM22

Phi-3-mini-4k-instruct Q4_K_M, llama_cpp 0.00.001.232 I srv llama_server: initializing ... defaults, -c 2048, 1x RTX 4090 (24GB): device memory 3.44 GiB (model 2229 + KV 768 + compute 62 MiB)3.4 GB3.7 GB

+7.8%

(1.08x)

通过

GM23

Phi-3-mini-4k-instruct Q4_K_M, llama_cpp 0.00.001.232 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 4.19 GiB (model 2229 + KV 1536 + compute 64 MiB)4.2 GB4.5 GB

+6.4%

(1.06x)

通过

GM24

Phi-3-mini-4k-instruct Q8_0, llama_cpp 0.00.000.784 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 5.7 GiB (model 3773 + KV 1536 + compute 64 MiB)5.7 GB5.9 GB

+4.2%

(1.04x)

通过

GM25

Phi-3-mini-4k-instruct (phi3:3.8b-mini-4k-instruct-q4_K_M), ollama ollama version is 0.34.4 defaults, default context (4096 for this VRAM tier), 1x RTX 4090 (24GB): device memory 4.19 GiB4.2 GB4.5 GB

+7.2%

(1.07x)

通过

GM26

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.965 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 8.93 GiB (model 4403 + KV 4096 + compute 140 MiB)8.9 GB9.0 GB

+0.9%

(1.01x)

通过

GM27

Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.965 I srv llama_server: initializing ... defaults, -c 131072, 1x A100-SXM4 (80GB): device memory 21.02 GiB (model 4403 + KV 16384 + compute 236 MiB)21.0 GB21.0 GB

-0.0%

(1.00x)

通过

GM28

Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.667 I srv llama_server: initializing ... defaults, -c 4096, 1x A100-SXM4 (80GB): device memory 19.77 GiB (model 18508 + KV 1024 + compute 196 MiB)19.8 GB19.3 GB

-2.5%

(0.97x)

通过

GM29

Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.667 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 26.79 GiB (model 18508 + KV 8192 + compute 224 MiB)26.8 GB26.3 GB

-1.9%

(0.98x)

通过

GM30

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.002.264 I srv llama_server: initializing ... defaults, -c 4096, 1x A100-SXM4 (80GB): device memory 4.91 GiB (model 4168 + KV 224 + compute 136 MiB)4.9 GB5.0 GB

+1.8%

(1.02x)

通过

GM31

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.002.264 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 6.88 GiB (model 4168 + KV 1792 + compute 164 MiB)6.9 GB6.5 GB

-5.1%

(0.95x)

通过

GM32

Qwen2.5-32B-Instruct (qwen2.5:32b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x A100-SXM4 (80GB): device memory 27.02 GiB27.0 GB26.5 GB

-1.8%

(0.98x)

通过

GM33

Qwen2.5-32B-Instruct (qwen2.5:32b), ollama ollama version is 0.34.4 defaults, num_ctx 8192, 1x A100-SXM4 (80GB): device memory 20.97 GiB21.0 GB20.5 GB

-2.1%

(0.98x)

通过

GM34

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x A100-SXM4 (80GB): device memory 6.64 GiB6.6 GB6.6 GB

-0.8%

(0.99x)

通过

GM35

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x A100-SXM4 (80GB): device memory 4.89 GiB4.9 GB5.1 GB

+3.5%

(1.03x)

通过

GM36

Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.557 I srv llama_server: initializing ... defaults, -c 8192, 1x H100 SXM (80GB): device memory 20.88 GiB (model 18508 + KV 2048 + compute 200 MiB)20.9 GB20.3 GB

-2.9%

(0.97x)

通过

GM37

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.683 I srv llama_server: initializing ... defaults, -c 4096, 1x H100 SXM (80GB): device memory 5.03 GiB (model 4168 + KV 224 + compute 136 MiB)5.0 GB5.0 GB

-0.6%

(0.99x)

通过

GM38

Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.683 I srv llama_server: initializing ... defaults, -c 32768, 1x H100 SXM (80GB): device memory 6.58 GiB (model 4168 + KV 1792 + compute 164 MiB)6.6 GB6.5 GB

-0.8%

(0.99x)

通过

GM39

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x H100 SXM (80GB): device memory 5.03 GiB5.0 GB5.1 GB

+0.6%

(1.01x)

通过

GM40

Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 32768, 1x H100 SXM (80GB): device memory 6.75 GiB6.8 GB6.6 GB

-2.4%

(0.98x)

通过

TM1

Qwen2.5-3B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (12733 x 32 tokens, free_gpu_memory_fraction 0.9)12733.0 GB12296.0 GB

-3.4%

(0.97x)

通过

TM2

Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (3458 x 32 tokens, free_gpu_memory_fraction 0.9)3458.0 GB3465.0 GB

+0.2%

(1.00x)

通过

TM3

Phi-3-mini-4k-instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (1116 x 32 tokens, free_gpu_memory_fraction 0.9)1116.0 GB1056.0 GB

-5.4%

(0.95x)

通过

TM4

Qwen2.5-14B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (7561 x 32 tokens, free_gpu_memory_fraction 0.9)7561.0 GB7550.0 GB

-0.1%

(1.00x)

通过

TM5

Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (32853 x 32 tokens, free_gpu_memory_fraction 0.9)32853.0 GB32872.0 GB

+0.1%

(1.00x)

通过

TM6

Llama-3.1-8B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (14132 x 32 tokens, free_gpu_memory_fraction 0.9)14132.0 GB14192.0 GB

+0.4%

(1.00x)

通过

TM7

Qwen2.5-14B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (7475 x 32 tokens, free_gpu_memory_fraction 0.9)7475.0 GB7550.0 GB

+1.0%

(1.01x)

通过

TM8

Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (32498 x 32 tokens, free_gpu_memory_fraction 0.9)32498.0 GB32872.0 GB

+1.2%

(1.01x)

通过

VM9

Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (65412 x 16 tokens, gpu_memory_utilization 0.92)65412.0 GB66519.0 GB

+1.7%

(1.02x)

通过