Llama-3-8B (8.03B) fp16 weights = safetensors size 16.1 GB 16.1 GB Pass
Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size 19.9 GB 19.5 GB Pass
Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size 26.9 GB 27.2 GB Pass
Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size 34.8 GB 33.3 GB Pass
Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size 47.4 GB 45.3 GB Pass
Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once 26.4 GB 27.6 GB Pass
Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size 14.5 GB 14.5 GB Pass
Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab) 6.4 GB 6.4 GB Pass
Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit) 4.7 GB 4.8 GB Pass
Llama-2-7B (6.74B) fp16 weights = safetensors size 13.5 GB 13.5 GB Pass
Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128) 1.1 GB 1.1 GB Pass
Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128) 4.3 GB 4.3 GB Pass
Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128) 1.3 GB 1.3 GB Pass
DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer) 2.3 GB 2.3 GB Pass
Kimi-K3 hybrid MLA + Linear Attention KV cache, 1M ctx, fp16 (93L, 576 cached dims/MLA layer, 74.2% linear reduction) 29.0 GB 29.0 GB Pass
Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128) 4.3 GB 4.3 GB Pass
Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128) 2.7 GB 2.7 GB Pass
Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128) 2.1 GB 2.1 GB Pass
Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128) 1.9 GB 1.9 GB Pass
Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress 17.2 GB 17.2 GB Pass
65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB) 44.0 GB 42.4 GB Pass
Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth) 8.5 GB 8.8 GB Pass
Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB) 48.0 GB 46.2 GB Pass
7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations) 112.0 GB 130.5 GB Pass
70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading 70.0 GB 45.3 GB Pass
Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100) 21.3 GB 22.1 GB Pass
End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded) 17.5 GB 17.5 GB Pass
End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB) 24.0 GB 24.8 GB Pass
Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight 405.0 GB 406.2 GB Pass
Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight 810.0 GB 810.0 GB Pass
Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress 33.8 GB 33.8 GB Pass
DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims) 10.3 GB 10.3 GB Pass
DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims) 4.7 GB 4.7 GB Pass
DeepSeek-V4.1-Flash KV cache, 1M ctx, fp16 (40L CED, CSA2 cross-layer reuse) 3.6 GB 3.6 GB Pass
DeepSeek-V4.1-Flash KV cache, 1M ctx, nvfp4 (40L CED, CSA2 890 B/tok) 0.9 GB 0.9 GB Pass
Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload) 54.7 GB 43.9 GB Pass
End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM) 486.8 GB 445.6 GB Pass
End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM) 946.2 GB 856.9 GB Pass
Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100) 14.2 GB 14.0 GB Pass
Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU) 36.7 GB 32.5 GB Pass
Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB) 12.5 GB 13.3 GB Pass
Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling 21.9 GB 17.3 GB Pass
Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB) 17.4 GB 20.9 GB Pass
Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB) 7.9 GB 10.5 GB Pass
Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU) 29.6 GB 25.4 GB Pass
Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU) 44.8 GB 49.4 GB Pass
Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB) 24.9 GB 23.9 GB Pass
33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B 21.0 GB 20.9 GB Pass
PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx 8.3 GB 7.0 GB Pass
PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx 31.6 GB 28.3 GB Pass
PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx 475.2 GB 342.8 GB Pass
PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx 215.4 GB 225.1 GB Pass
Qwen3-Omni-30B-A3B (33.0B total) BF16 weights = safetensors size 66.0 GB 66.0 GB Pass
Qwen3-Omni-30B-A3B (33.0B total, 3B aux) W4A16 weights = AutoRound int4 size 25.0 GB 23.8 GB Pass
End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE) 2100.0 GB 2073.6 GB Pass
Llama-3.1-70B KV Cache with batch size 32 85.9 GB 85.9 GB Pass
DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes) 76.5 GB 81.0 GB Pass
Inkling-Small (276.0B / 12.0B active) BF16 weights = base safetensors size 552.0 GB 552.0 GB Pass
Inkling-Small (276.0B / 12.0B active) NVFP4 weights 140.0 GB 140.5 GB Pass
Nemotron 3.5 Lightning (30.0B / 3.0B active) BF16 weights = safetensors size 60.0 GB 60.0 GB Pass
Nemotron 3.5 Lightning (30.0B / 3.0B active) NVFP4 weights 16.0 GB 16.1 GB Pass
Qwen3.8-2.4T-A95B (2400.0B / 95.0B active) NVFP4 weights 1200.0 GB 1206.1 GB Pass
Inkling-Small KV cache, 8k ctx, fp16 (42L, 8 KV heads, head 128) 1.4 GB 1.4 GB Pass
Nemotron 3.5 Lightning KV cache, 8k ctx, fp16 (52L, hidden 2688, linear attention ratio 0.75) 1.1 GB 1.1 GB Pass
Muse-Glimmer-30B (29.6B) bf16 weights = safetensors size 59.2 GB 59.2 GB Pass
Muse-Glimmer-30B KV cache, 8k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048) 0.2 GB 0.2 GB Pass
Muse-Glimmer-30B KV cache, 128k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048) 1.8 GB 1.8 GB Pass
End-to-end: Muse-Glimmer-30B Q4_K_M inference, 8k ctx (claimed ~17-24GB on RTX 4090/5090) 19.5 GB 22.2 GB Pass
Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 1 (vLLM memory profiling: ~140GB weights + ~42GB KV + ~19GB framework/NCCL) 203.0 GB 203.6 GB Pass
Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 4 (vLLM memory profiling with chunked prefill: ~205GB total across cluster) 205.5 GB 205.8 GB Pass
Llama-3.2-11B-Vision (10.67B total, 2.64B aux) BF16 weights = safetensors size 21.3 GB 21.3 GB Pass
Qwen2-VL-7B-Instruct (8.29B untied safetensors) BF16 weights = safetensors size 16.6 GB 16.6 GB Pass
Qwen2-VL-7B-Instruct-GPTQ-Int4 weights = GPTQ Int4 safetensors size 6.9 GB 6.8 GB Pass
Qwen2.5-3B-Instruct (3.09B) BF16 weights = published safetensors size 6.2 GB 6.2 GB Pass
Mistral-7B-Instruct-v0.3 (7.25B) BF16 weights = published safetensors size 14.5 GB 14.5 GB Pass
Qwen2.5-0.5B-Instruct (0.49B) BF16 weights = published single safetensors size 1.0 GB 1.0 GB Pass
Qwen3.5-0.8B (0.87B total with 0.10B visual encoder) BF16 weights = published safetensors size 1.8 GB 1.7 GB Pass
Gemma 4 E2B (5.12B total with 0.49B vision+audio) BF16 weights = published safetensors size 10.3 GB 10.2 GB Pass
Qwen2.5-3B KV cache, 8k ctx, fp16 (36L, 2 KV heads, head 128) 0.3 GB 0.3 GB Pass
Qwen2.5-3B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 8.01GB) 8.0 GB 8.1 GB Pass
Qwen2.5-3B-Instruct FP16 raw unchunked prefill, RTX 4090 (24GB), 8k ctx, batch 4 -> OOM Insufficient Insufficient - Pass
DeepSeek-V2-Lite (16B MoE total / 2.4B active) FP16 weights loading on RTX 4090 (24GB) -> OOM Insufficient Insufficient - Pass
Qwen2-VL-2B (2.21B total with auxiliary visual encoder) BF16 empirical weights on RTX 4090 4.4 GB 4.4 GB Pass
Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 15.41GB) 15.4 GB 16.5 GB Pass
Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 4 (measured peak: 18.59GB) 18.6 GB 19.3 GB Pass
Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 8 (measured peak: 23.05GB) 23.1 GB 23.6 GB Pass
Qwen2.5-7B-Instruct FP16 empirical model weights on RTX 4090 15.2 GB 15.2 GB Pass
Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 1k ctx, batch 1 (measured peak: 18.40GB) 18.4 GB 19.9 GB Pass
Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 19.87GB) 19.9 GB 20.3 GB Pass
Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 1 (measured peak: 17.19GB) 17.2 GB 17.7 GB Pass
Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 8 (measured peak: 32.73GB) 32.7 GB 28.2 GB Pass
Llama-3.1-8B-Instruct QLoRA r=16 training, 1x RTX 4090 (24GB), bs=2 seq=2048 (measured peak: 19.71GB) 19.7 GB 15.0 GB Pass
Qwen2.5-3B-Instruct full FT, 1x A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 33.59GB) 33.6 GB 32.6 GB Pass
DeepSeek-R1-Distill-Qwen-14B FP16 empirical model weights, A100-SXM4 (80GB) 29.5 GB 29.5 GB Pass
DeepSeek-R1-Distill-Qwen-14B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 30.30GB) 30.3 GB 31.1 GB Pass
DeepSeek-R1-Distill-Qwen-7B FP16 empirical model weights, RTX 4090 (24GB) 15.2 GB 15.2 GB Pass
DeepSeek-R1-Distill-Qwen-7B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 16.23GB) 16.2 GB 16.5 GB Pass
DeepSeek-R1-Distill-Qwen-32B FP16 empirical model weights, A100-SXM4 (80GB) 65.5 GB 65.5 GB Pass
DeepSeek-R1-Distill-Qwen-32B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 66.65GB) 66.7 GB 67.3 GB Pass
Qwen2.5-3B-Instruct FP16 inference, H200 (141GB), 2k ctx, batch 1 (measured peak: 6.76GB) 6.8 GB 7.3 GB Pass
Qwen3.8-27B FP16 empirical model weights, A100-SXM4 (80GB) 53.8 GB 53.8 GB Pass
Qwen3.8-27B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 54.67GB) 54.7 GB 55.3 GB Pass
Ministral-3-8B-Instruct-2512 FP16 empirical model weights, A100-SXM4 (80GB) 17.8 GB 17.8 GB Pass
Ministral-3-8B-Instruct-2512 FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 20.38GB) 20.4 GB 19.3 GB Pass
Granite-4.2-8B FP16 empirical model weights, RTX 4090 (24GB) 17.6 GB 17.6 GB Pass
Granite-4.2-8B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 18.89GB) 18.9 GB 19.1 GB Pass
Gemma-4-E2B FP16 empirical model weights, RTX 4090 (24GB) 10.2 GB 10.2 GB Pass
Gemma-4-E2B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 12.81GB) 12.8 GB 13.1 GB Pass
Gemma-4-12B FP16 empirical model weights, L40S (48GB) 23.9 GB 23.9 GB Pass
Gemma-4-12B FP16 inference, L40S (48GB), 2k ctx, batch 1 (measured peak: 26.62GB) 26.6 GB 26.6 GB Pass
Ministral-3-3B Full FT 32-bit AdamW, A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 36.95GB) 37.0 GB 71.5 GB Miss
DeepSeek-R1-Distill-Qwen-7B LoRA 32-bit AdamW training, RTX 4090 (24GB) -> OOM Insufficient Insufficient - Pass
Ministral-3-8B-Instruct-2512 LoRA 8-bit AdamW training, RTX 4090 (24GB) -> OOM Insufficient Insufficient - Pass
Qwen3.8-Flash-Next (176.9B) FP16 inference, 2x H200 (141GB) -> OOM (weights exceed cluster capacity) Insufficient Insufficient - Pass