趋近智
Predictions and speed estimates compared against verified hardware benchmarks and ground-truth configurations.
校准报告
| 测试套件 | 测试用例数 | 带内通过率 | 中位数误差 | GMFE 因子误差 | 系统偏差 | 最差情况 |
|---|---|---|---|---|---|---|
| 显存占用 | 56 | 100.0% | 3.5% | 1.07x | 微偏低 | D5 (0.65x) |
| 解码速度 | 47 | 100.0% | 16.5% | 1.25x | 偏低 | T9 (0.30x) |
| 微调速度 | 18 | 100.0% | 22.0% | 1.33x | 偏低 | T2 (0.46x) |
内存预测具有高度确定性且非常准确,而速度估算对运行和硬件配置更为敏感。未命中主要是由于源基准中未指定参数数据所致。
带内通过率
100.0%
预测符合目标误差范围的比例
100%
中位数误差
3.5%
典型的绝对预测误差
GMFE 因子误差
1.07x
几何平均误差因子
系统偏差
微偏低
整体校准偏差方向
| ID | 场景与配置 | 测量值 | 预测值 | 比例 | 状态 |
|---|---|---|---|---|---|
A1 | Llama-3-8B (8.03B) fp16 weights = safetensors size | 16.1 GB | 16.1 GB | 1.00x | 通过 |
A2 | Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size | 19.9 GB | 19.5 GB | 0.98x | 通过 |
A3 | Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size | 26.9 GB | 27.2 GB | 1.01x | 通过 |
A4 | Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size | 34.8 GB | 33.3 GB | 0.96x | 通过 |
A5 | Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size | 47.4 GB | 45.3 GB | 0.96x | 通过 |
A6 | Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once | 26.4 GB | 27.6 GB | 1.04x | 通过 |
A7 | Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size | 14.5 GB | 14.5 GB | 1.00x | 通过 |
A8 | Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab) | 6.4 GB | 6.4 GB | 1.00x | 通过 |
A9 | Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit) | 4.7 GB | 4.8 GB | 1.02x | 通过 |
A10 | Llama-2-7B (6.74B) fp16 weights = safetensors size | 13.5 GB | 13.5 GB | 1.00x | 通过 |
B1 | Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128) | 1.1 GB | 1.1 GB | 1.00x | 通过 |
B2 | Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128) | 4.3 GB | 4.3 GB | 1.00x | 通过 |
B3 | Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128) | 1.3 GB | 1.3 GB | 1.00x | 通过 |
B4 | DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer) | 2.3 GB | 2.3 GB | 1.00x | 通过 |
B5 | Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128) | 4.3 GB | 4.3 GB | 1.00x | 通过 |
B6 | Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128) | 2.7 GB | 2.7 GB | 1.00x | 通过 |
B7 | Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128) | 2.1 GB | 2.1 GB | 1.00x | 通过 |
B8 | Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128) | 1.9 GB | 1.9 GB | 1.00x | 通过 |
B9 | Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress | 17.2 GB | 17.2 GB | 1.00x | 通过 |
D1 | 65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB) | 44.0 GB | 42.4 GB | 0.96x | 通过 |
D2 | Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth) | 8.5 GB | 8.8 GB | 1.03x | 通过 |
D3 | Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB) | 48.0 GB | 46.2 GB | 0.96x | 通过 |
D4 | 7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations) | 112.0 GB | 130.5 GB | 1.17x | 通过 |
D5 | 70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading | 70.0 GB | 45.3 GB | 0.65x | 通过 |
D7 | Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100) | 21.3 GB | 22.1 GB | 1.04x | 通过 |
E1 | End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded) | 17.5 GB | 17.5 GB | 1.00x | 通过 |
E2 | End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB) | 24.0 GB | 24.8 GB | 1.03x | 通过 |
A11 | Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight | 405.0 GB | 406.2 GB | 1.00x | 通过 |
A12 | Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight | 810.0 GB | 810.0 GB | 1.00x | 通过 |
B10 | Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress | 33.8 GB | 33.8 GB | 1.00x | 通过 |
B11 | DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims) | 10.3 GB | 10.3 GB | 1.00x | 通过 |
B12 | DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims) | 4.7 GB | 4.7 GB | 1.00x | 通过 |
D8 | Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload) | 54.7 GB | 43.9 GB | 0.80x | 通过 |
E3 | End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM) | 486.8 GB | 445.6 GB | 0.92x | 通过 |
E4 | End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM) | 946.2 GB | 856.9 GB | 0.91x | 通过 |
D9 | Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100) | 14.2 GB | 13.6 GB | 0.96x | 通过 |
D10 | Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU) | 36.7 GB | 32.5 GB | 0.89x | 通过 |
D11 | Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB) | 12.5 GB | 13.3 GB | 1.06x | 通过 |
D12 | Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling | 21.9 GB | 17.3 GB | 0.79x | 通过 |
D13 | Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB) | 17.4 GB | 20.9 GB | 1.20x | 通过 |
D14 | Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB) | 7.9 GB | 10.5 GB | 1.33x | 通过 |
D15 | Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU) | 29.6 GB | 25.4 GB | 0.86x | 通过 |
D16 | Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU) | 44.8 GB | 49.4 GB | 1.10x | 通过 |
D17 | Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB) | 24.9 GB | 23.9 GB | 0.96x | 通过 |
D18 | 33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B | 21.0 GB | 20.9 GB | 0.99x | 通过 |
E5 | PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx | 8.3 GB | 7.0 GB | 0.84x | 通过 |
E6 | PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx | 31.6 GB | 28.3 GB | 0.89x | 通过 |
E7 | PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx | 475.2 GB | 351.1 GB | 0.74x | 通过 |
E8 | PK 44: DeepSeek-R1 671B Q8, FP8 KV, 8x B200 (192GB), 131k ctx | 1056.1 GB | 883.5 GB | 0.84x | 通过 |
E9 | PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx | 215.4 GB | 239.6 GB | 1.11x | 通过 |
A13 | Qwen3-Omni-30B-A3B BF16 weights = safetensors size | 66.0 GB | 69.0 GB | 1.05x | 通过 |
A14 | Qwen3-Omni-30B-A3B W4A16 weights = AutoRound int4 size | 25.0 GB | 20.5 GB | 0.82x | 通过 |
E10 | End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE) | 2100.0 GB | 2073.6 GB | 0.99x | 通过 |
B10 | Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress | 85.9 GB | 85.9 GB | 1.00x | 通过 |
D19 | DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes) | 76.5 GB | 81.0 GB | 1.06x | 通过 |
E11 | Llama-3.1-405B FP8 Inference on 8x H100 (TP=8) with batch size 64 | 730.0 GB | 730.0 GB | 1.00x | 通过 |
APX AI
在线
我可以读取您正在浏览的页面。随时向我提问!