ApX 标志ApX 标志

趋近智

精度评分卡

Predictions and speed estimates compared against verified hardware benchmarks and ground-truth configurations.

Go To Calculator

校准报告

测试套件测试用例数带内通过率中位数误差GMFE 因子误差系统偏差最差情况
显存占用56100.0%3.5%1.07x

微偏低

D5 (0.65x)
解码速度47100.0%16.5%1.25x

偏低

T9 (0.30x)
微调速度18100.0%22.0%1.33x

偏低

T2 (0.46x)

内存预测具有高度确定性且非常准确,而速度估算对运行和硬件配置更为敏感。未命中主要是由于源基准中未指定参数数据所致。

带内通过率

100.0%

预测符合目标误差范围的比例

100%

中位数误差

3.5%

典型的绝对预测误差

GMFE 因子误差

1.07x

几何平均误差因子

系统偏差

微偏低

整体校准偏差方向

ID场景与配置测量值预测值比例状态

A1

Llama-3-8B (8.03B) fp16 weights = safetensors size16.1 GB16.1 GB1.00x

通过

A2

Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size19.9 GB19.5 GB0.98x

通过

A3

Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size26.9 GB27.2 GB1.01x

通过

A4

Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size34.8 GB33.3 GB0.96x

通过

A5

Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size47.4 GB45.3 GB0.96x

通过

A6

Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once26.4 GB27.6 GB1.04x

通过

A7

Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size14.5 GB14.5 GB1.00x

通过

A8

Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab)6.4 GB6.4 GB1.00x

通过

A9

Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit)4.7 GB4.8 GB1.02x

通过

A10

Llama-2-7B (6.74B) fp16 weights = safetensors size13.5 GB13.5 GB1.00x

通过

B1

Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128)1.1 GB1.1 GB1.00x

通过

B2

Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128)4.3 GB4.3 GB1.00x

通过

B3

Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128)1.3 GB1.3 GB1.00x

通过

B4

DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer)2.3 GB2.3 GB1.00x

通过

B5

Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128)4.3 GB4.3 GB1.00x

通过

B6

Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128)2.7 GB2.7 GB1.00x

通过

B7

Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128)2.1 GB2.1 GB1.00x

通过

B8

Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128)1.9 GB1.9 GB1.00x

通过

B9

Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress17.2 GB17.2 GB1.00x

通过

D1

65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB)44.0 GB42.4 GB0.96x

通过

D2

Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth)8.5 GB8.8 GB1.03x

通过

D3

Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB)48.0 GB46.2 GB0.96x

通过

D4

7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations)112.0 GB130.5 GB1.17x

通过

D5

70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading70.0 GB45.3 GB0.65x

通过

D7

Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100)21.3 GB22.1 GB1.04x

通过

E1

End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded)17.5 GB17.5 GB1.00x

通过

E2

End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB)24.0 GB24.8 GB1.03x

通过

A11

Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight405.0 GB406.2 GB1.00x

通过

A12

Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight810.0 GB810.0 GB1.00x

通过

B10

Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress33.8 GB33.8 GB1.00x

通过

B11

DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims)10.3 GB10.3 GB1.00x

通过

B12

DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims)4.7 GB4.7 GB1.00x

通过

D8

Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload)54.7 GB43.9 GB0.80x

通过

E3

End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM)486.8 GB445.6 GB0.92x

通过

E4

End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM)946.2 GB856.9 GB0.91x

通过

D9

Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100)14.2 GB13.6 GB0.96x

通过

D10

Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU)36.7 GB32.5 GB0.89x

通过

D11

Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB)12.5 GB13.3 GB1.06x

通过

D12

Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling21.9 GB17.3 GB0.79x

通过

D13

Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB)17.4 GB20.9 GB1.20x

通过

D14

Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB)7.9 GB10.5 GB1.33x

通过

D15

Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU)29.6 GB25.4 GB0.86x

通过

D16

Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU)44.8 GB49.4 GB1.10x

通过

D17

Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB)24.9 GB23.9 GB0.96x

通过

D18

33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B21.0 GB20.9 GB0.99x

通过

E5

PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx8.3 GB7.0 GB0.84x

通过

E6

PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx31.6 GB28.3 GB0.89x

通过

E7

PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx475.2 GB351.1 GB0.74x

通过

E8

PK 44: DeepSeek-R1 671B Q8, FP8 KV, 8x B200 (192GB), 131k ctx1056.1 GB883.5 GB0.84x

通过

E9

PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx215.4 GB239.6 GB1.11x

通过

A13

Qwen3-Omni-30B-A3B BF16 weights = safetensors size66.0 GB69.0 GB1.05x

通过

A14

Qwen3-Omni-30B-A3B W4A16 weights = AutoRound int4 size25.0 GB20.5 GB0.82x

通过

E10

End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE)2100.0 GB2073.6 GB0.99x

通过

B10

Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress85.9 GB85.9 GB1.00x

通过

D19

DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes)76.5 GB81.0 GB1.06x

通过

E11

Llama-3.1-405B FP8 Inference on 8x H100 (TP=8) with batch size 64730.0 GB730.0 GB1.00x

通过