Predictions and speed estimates compared against verified hardware benchmarks and ground-truth configurations.
Calibration Report
| Evaluation Suite | Scored Cases | In-Band Pass Rate | Median Error | GMFE | Systemic Bias | Worst Case |
|---|---|---|---|---|---|---|
| VRAM Usage | 56 | 100.0% | 3.5% | 1.07x | Slight Underpredict | D5 (0.65x) |
| Decode Speed | 47 | 100.0% | 16.5% | 1.25x | Underpredicts | T9 (0.30x) |
| Training Speed | 18 | 100.0% | 22.0% | 1.33x | Underpredicts | T2 (0.46x) |
Memory predictions are highly deterministic and very accurate, while speed estimates are more sensitive to runtime and hardware configurations. Misses are largely due to unspecified parameter data in the source benchmarks.
In-Band Pass Rate
100.0%
Predictions matching target boundaries
100%
Median Error
3.5%
Typical absolute prediction error
GMFE
1.07x
Geometric mean error factor
Systemic Bias
Slight Underpredict
Overall calibration offset direction
| ID | Scenario & Config | Measured | Predicted | Ratio | Status |
|---|---|---|---|---|---|
A1 | Llama-3-8B (8.03B) fp16 weights = safetensors size | 16.1 GB | 16.1 GB | 1.00x | Pass |
A2 | Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size | 19.9 GB | 19.5 GB | 0.98x | Pass |
A3 | Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size | 26.9 GB | 27.2 GB | 1.01x | Pass |
A4 | Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size | 34.8 GB | 33.3 GB | 0.96x | Pass |
A5 | Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size | 47.4 GB | 45.3 GB | 0.96x | Pass |
A6 | Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once | 26.4 GB | 27.6 GB | 1.04x | Pass |
A7 | Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size | 14.5 GB | 14.5 GB | 1.00x | Pass |
A8 | Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab) | 6.4 GB | 6.4 GB | 1.00x | Pass |
A9 | Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit) | 4.7 GB | 4.8 GB | 1.02x | Pass |
A10 | Llama-2-7B (6.74B) fp16 weights = safetensors size | 13.5 GB | 13.5 GB | 1.00x | Pass |
B1 | Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128) | 1.1 GB | 1.1 GB | 1.00x | Pass |
B2 | Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128) | 4.3 GB | 4.3 GB | 1.00x | Pass |
B3 | Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128) | 1.3 GB | 1.3 GB | 1.00x | Pass |
B4 | DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer) | 2.3 GB | 2.3 GB | 1.00x | Pass |
B5 | Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128) | 4.3 GB | 4.3 GB | 1.00x | Pass |
B6 | Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128) | 2.7 GB | 2.7 GB | 1.00x | Pass |
B7 | Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128) | 2.1 GB | 2.1 GB | 1.00x | Pass |
B8 | Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128) | 1.9 GB | 1.9 GB | 1.00x | Pass |
B9 | Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress | 17.2 GB | 17.2 GB | 1.00x | Pass |
D1 | 65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB) | 44.0 GB | 42.4 GB | 0.96x | Pass |
D2 | Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth) | 8.5 GB | 8.8 GB | 1.03x | Pass |
D3 | Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB) | 48.0 GB | 46.2 GB | 0.96x | Pass |
D4 | 7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations) | 112.0 GB | 130.5 GB | 1.17x | Pass |
D5 | 70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading | 70.0 GB | 45.3 GB | 0.65x | Pass |
D7 | Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100) | 21.3 GB | 22.1 GB | 1.04x | Pass |
E1 | End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded) | 17.5 GB | 17.5 GB | 1.00x | Pass |
E2 | End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB) | 24.0 GB | 24.8 GB | 1.03x | Pass |
A11 | Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight | 405.0 GB | 406.2 GB | 1.00x | Pass |
A12 | Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight | 810.0 GB | 810.0 GB | 1.00x | Pass |
B10 | Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress | 33.8 GB | 33.8 GB | 1.00x | Pass |
B11 | DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims) | 10.3 GB | 10.3 GB | 1.00x | Pass |
B12 | DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims) | 4.7 GB | 4.7 GB | 1.00x | Pass |
D8 | Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload) | 54.7 GB | 43.9 GB | 0.80x | Pass |
E3 | End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM) | 486.8 GB | 445.6 GB | 0.92x | Pass |
E4 | End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM) | 946.2 GB | 856.9 GB | 0.91x | Pass |
D9 | Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100) | 14.2 GB | 13.6 GB | 0.96x | Pass |
D10 | Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU) | 36.7 GB | 32.5 GB | 0.89x | Pass |
D11 | Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB) | 12.5 GB | 13.3 GB | 1.06x | Pass |
D12 | Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling | 21.9 GB | 17.3 GB | 0.79x | Pass |
D13 | Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB) | 17.4 GB | 20.9 GB | 1.20x | Pass |
D14 | Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB) | 7.9 GB | 10.5 GB | 1.33x | Pass |
D15 | Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU) | 29.6 GB | 25.4 GB | 0.86x | Pass |
D16 | Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU) | 44.8 GB | 49.4 GB | 1.10x | Pass |
D17 | Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB) | 24.9 GB | 23.9 GB | 0.96x | Pass |
D18 | 33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B | 21.0 GB | 20.9 GB | 0.99x | Pass |
E5 | PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx | 8.3 GB | 7.0 GB | 0.84x | Pass |
E6 | PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx | 31.6 GB | 28.3 GB | 0.89x | Pass |
E7 | PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx | 475.2 GB | 351.1 GB | 0.74x | Pass |
E8 | PK 44: DeepSeek-R1 671B Q8, FP8 KV, 8x B200 (192GB), 131k ctx | 1056.1 GB | 883.5 GB | 0.84x | Pass |
E9 | PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx | 215.4 GB | 239.6 GB | 1.11x | Pass |
A13 | Qwen3-Omni-30B-A3B BF16 weights = safetensors size | 66.0 GB | 69.0 GB | 1.05x | Pass |
A14 | Qwen3-Omni-30B-A3B W4A16 weights = AutoRound int4 size | 25.0 GB | 20.5 GB | 0.82x | Pass |
E10 | End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE) | 2100.0 GB | 2073.6 GB | 0.99x | Pass |
B10 | Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress | 85.9 GB | 85.9 GB | 1.00x | Pass |
D19 | DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes) | 76.5 GB | 81.0 GB | 1.06x | Pass |
E11 | Llama-3.1-405B FP8 Inference on 8x H100 (TP=8) with batch size 64 | 730.0 GB | 730.0 GB | 1.00x | Pass |
©2025 ApX Machine Learning
APX AI
Online
I can see the page you're looking at. Ask me anything!