Llama-3-8B (8.03B) fp16 weights = safetensors size 15.0 GB 15.0 GB Pass
Qwen2.5-32B (32.8B) Q4_K_M weights = GGUF file size 18.5 GB 18.2 GB Pass
Qwen2.5-32B (32.8B) Q6_K weights = GGUF file size 25.1 GB 25.3 GB Pass
Qwen2.5-32B (32.8B) Q8_0 weights = GGUF file size 32.4 GB 31.0 GB Pass
Qwen2.5-72B (72.7B) Q4_K_M weights = GGUF file size 44.1 GB 42.2 GB Pass
Mixtral 8x7B MoE (46.7B total / 12.9B active) Q4_K_M weights, counted once 24.6 GB 25.7 GB Pass
Mistral-7B-Instruct-v0.3 (7.25B) bf16 weights = safetensors size 13.5 GB 13.5 GB Pass
Llama-3.2-3B (3.21B) bf16 weights = safetensors size (small model, large vocab) 6.0 GB 6.0 GB Pass
Qwen2.5-7B (7.62B) Q4_K_M weights = GGUF file size (large-vocab model at 4-bit) 4.4 GB 4.4 GB Pass
Llama-2-7B (6.74B) fp16 weights = safetensors size 12.6 GB 12.6 GB Pass
Llama-3-8B KV cache, 8k ctx, fp16 (32L, 8 KV heads, head 128) 1.0 GB 1.0 GB Pass
Qwen2.5-32B KV cache, 16k ctx, fp16 (64L, 8 KV heads, head 128) 4.0 GB 4.0 GB Pass
Llama-3-70B KV cache, 4k ctx, fp16 (80L, 8 KV heads, head 128) 1.3 GB 1.3 GB Pass
DeepSeek-V3 MLA KV cache, 32k ctx, bf16 (61L, 576 cached dims/token/layer) 2.1 GB 2.1 GB Pass
Kimi-K3 hybrid MLA + Linear Attention KV cache, 1M ctx, fp16 (93L, 576 cached dims/MLA layer, 74.2% linear reduction) 27.0 GB 27.0 GB Pass
Mistral-7B-v0.3 KV cache, 32k ctx, fp16 (32L, 8 KV heads, head 128) 4.0 GB 4.0 GB Pass
Qwen2.5-72B KV cache, 8k ctx, fp16 (80L, 8 KV heads, head 128) 2.5 GB 2.5 GB Pass
Llama-2-7B MHA KV cache, 4k ctx, fp16 (32L, 32 KV heads, head 128) 2.0 GB 2.0 GB Pass
Qwen2.5-7B KV cache, 32k ctx, fp16 (28L, 4 KV heads, head 128) 1.8 GB 1.8 GB Pass
Llama-3.1-8B KV cache, 128k ctx, fp16 (32L, 8 KV heads, head 128) - long-context stress 16.0 GB 16.0 GB Pass
65B QLoRA (nf4, paged optimizer), 1x RTX A6000 (48GB) 44.0 GB 42.9 GB Pass
Llama-3.1-8B QLoRA rank 32, batch 1, 2k ctx (~9GB HF+FA2, <8GB Unsloth) 8.5 GB 7.4 GB Miss
Llama-3-70B QLoRA (~48GB, fits 1x A100 80GB) 48.0 GB 42.7 GB Miss
7B full FT, mixed-precision AdamW (~16 bytes/param static = 112GB before activations) 104.3 GB 122.3 GB Pass
70B full FT, 8x H100 (80GB), ZeRO-3, gradient checkpointing, CPU offloading 70.0 GB 47.0 GB Pass
Llama-2-7B 16-bit LoRA r=8 (q+v only), micro-batch 1 (measured 21.33GB on A100) 21.3 GB 21.2 GB Pass
End-to-end: Llama-3-8B fp16 inference, 2k ctx (measured ~17.5GB loaded) 17.5 GB 16.3 GB Miss
End-to-end: Qwen2.5-32B Q4_K_M inference, 16k ctx (measured ~24GB) 24.0 GB 23.2 GB Pass
Llama-3.1-405B FP8 weights = 1.0 byte per parameter base weight 377.2 GB 378.3 GB Pass
Llama-3.1-405B FP16 weights = 2.0 bytes per parameter base weight 754.4 GB 754.4 GB Pass
Llama-3.1-405B KV cache, 131072 ctx, fp8 (126L, 8 KV heads, head 128) - cluster context stress 31.5 GB 31.5 GB Pass
DeepSeek-V4-Pro KV cache, 1M ctx, bf16 (61L, 512 shared / 128 indexer dims) 9.6 GB 9.6 GB Pass
DeepSeek-V4-Pro KV cache, 1M ctx, fp8/fp4 (61L, 512 shared / 128 indexer dims) 4.3 GB 4.3 GB Pass
DeepSeek-V4.1-Flash KV cache, 1M ctx, fp16 (40L CED, CSA2 cross-layer reuse) 3.4 GB 3.4 GB Pass
DeepSeek-V4.1-Flash KV cache, 1M ctx, nvfp4 (40L CED, CSA2 890 B/tok) 0.9 GB 0.9 GB Pass
Llama-3-70B full FT, 32x H100 80GB, ZeRO-3 (no offload) 54.7 GB 45.7 GB Pass
End-to-end: Llama-3.1-405B FP8 inference, 8x H100 (80GB), 2k ctx (fits in ~486.8GB VRAM) 486.8 GB 417.2 GB Miss
End-to-end: Llama-3.1-405B FP16 inference, 8x H200 (141GB), 2k ctx (fits in ~946.21GB VRAM) 881.2 GB 800.3 GB Pass
Llama-2-7B QLoRA NF4 r=8 (q+v), micro-batch 1 (measured 14.18GB on A100) 14.2 GB 13.8 GB Pass
Llama-2-7B full FT bf16 mixed, sharded on 2x A100 (measured 36.66GB/GPU) 36.7 GB 31.0 GB Pass
Mistral-7B QLoRA r=16, batch 4, 2k ctx, A100-40 (HF+FA2 19.4GB, Unsloth 10.3GB) 12.5 GB 13.1 GB Pass
Gemma-7B QLoRA, batch 1, 8k ctx, A100-80 (Unsloth 21.9GB, HF+FA2 47.8GB) - long-seq activation scaling 21.9 GB 16.9 GB Pass
Llama-3.1-8B LoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 16.2 GiB = 17.4GB) 17.4 GB 20.9 GB Pass
Llama-3.1-8B QLoRA r=8, batch 2, packed 2k, RTX 4090 (torchtune 7.4 GiB = 7.9GB) 7.9 GB 10.3 GB Pass
Llama-3.1-70B LoRA, 8x A100 FSDP full-shard (torchtune 27.6 GiB = 29.6GB/GPU) 29.6 GB 24.9 GB Pass
Llama-3.1-405B QLoRA, 8x A100 FSDP (torchtune 44.8GB/GPU) 44.8 GB 45.3 GB Pass
Llama-3.1-8B LoRA r=16 (q+v), batch 2, seq 512, NO checkpointing/FA (measured ~24.9GB, RTX PRO 5000 48GB) 24.9 GB 22.9 GB Pass
33B QLoRA (nf4, paged optimizer), 1x RTX 4090 (24GB) - Guanaco-33B 21.0 GB 21.7 GB Pass
PK 35: Qwen3.5-9B Q4_K_M, INT8 KV, RTX 5070 Ti (16GB), 8k ctx 8.3 GB 6.6 GB Miss
PK 37: Qwen3.5-35B-A3B nvfp6, RTX 5090 (32GB), 1k ctx 31.6 GB 26.4 GB Miss
PK 38: Qwen3.5-397B-A17B nvfp4, 8x RTX 6000 Blackwell (96GB), 65k ctx 475.2 GB 319.4 GB Miss
PK 58: Qwen3.5-122B-A10B Q8, 8x RTX 5000 Blackwell (48GB), 20k ctx 215.4 GB 209.8 GB Pass
Qwen3-Omni-30B-A3B (33.0B total) BF16 weights = safetensors size 61.5 GB 61.5 GB Pass
Qwen3-Omni-30B-A3B (33.0B total, 3B aux) W4A16 weights = AutoRound int4 size 23.3 GB 22.2 GB Pass
End-to-end: Kimi K3 MXFP4 on 8x B300 (2.8T total parameters, 16/896 MoE) 2100.0 GB 2073.6 GB Pass
Llama-3.1-70B KV Cache with batch size 32 80.0 GB 80.0 GB Pass
DeepSpeed ZeRO-3 Full Finetuning of a 70B model with batch size 16 on 16x H100 (2 nodes) 76.5 GB 80.3 GB Pass
Inkling-Small (276.0B / 12.0B active) BF16 weights = base safetensors size 514.1 GB 514.1 GB Pass
Inkling-Small (276.0B / 12.0B active) NVFP4 weights 130.4 GB 130.8 GB Pass
Nemotron 3.5 Lightning (30.0B / 3.0B active) BF16 weights = safetensors size 55.9 GB 55.9 GB Pass
Nemotron 3.5 Lightning (30.0B / 3.0B active) NVFP4 weights 14.9 GB 14.9 GB Pass
Qwen3.8-2.4T-A95B (2400.0B / 95.0B active) NVFP4 weights 1117.6 GB 1123.3 GB Pass
Inkling-Small KV cache, 8k ctx, fp16 (42L, 8 KV heads, head 128) 1.3 GB 1.3 GB Pass
Nemotron 3.5 Lightning KV cache, 8k ctx, fp16 (52L, hidden 2688, linear attention ratio 0.75) 1.1 GB 1.1 GB Pass
Muse-Glimmer-30B (29.6B) bf16 weights = safetensors size 55.1 GB 55.1 GB Pass
Muse-Glimmer-30B KV cache, 8k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048) 0.2 GB 0.2 GB Pass
Muse-Glimmer-30B KV cache, 128k ctx, fp16 (52L, 2 KV heads, head 128, 0.75 sliding window 2048) 1.7 GB 1.7 GB Pass
End-to-end: Muse-Glimmer-30B Q4_K_M inference, 8k ctx (claimed ~17-24GB on RTX 4090/5090) 19.5 GB 20.8 GB Pass
Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 1 (vLLM memory profiling: ~140GB weights + ~42GB KV + ~19GB framework/NCCL) 203.0 GB 205.5 GB Pass
Llama-3.1-70B FP16 inference, 8x H100 (80GB), TP=8, 128k ctx, batch 4 (vLLM memory profiling with chunked prefill: ~205GB total across cluster) 205.5 GB 205.5 GB Pass
Llama-3.2-11B-Vision (10.67B total, 2.64B aux) BF16 weights = safetensors size 19.9 GB 19.9 GB Pass
Qwen2-VL-7B-Instruct (8.29B untied safetensors) BF16 weights = safetensors size 15.4 GB 15.4 GB Pass
Qwen2-VL-7B-Instruct-GPTQ-Int4 weights = GPTQ Int4 safetensors size 6.5 GB 6.3 GB Pass
Qwen2.5-3B-Instruct (3.09B) BF16 weights = published safetensors size 5.8 GB 5.8 GB Pass
Mistral-7B-Instruct-v0.3 (7.25B) BF16 weights = published safetensors size 13.5 GB 13.5 GB Pass
Qwen2.5-0.5B-Instruct (0.49B) BF16 weights = published single safetensors size 0.9 GB 0.9 GB Pass
Qwen3.5-0.8B (0.87B total with 0.10B visual encoder) BF16 weights = published safetensors size 1.6 GB 1.6 GB Pass
Gemma 4 E2B (5.12B total with 0.49B vision+audio) BF16 weights = published safetensors size 9.6 GB 9.5 GB Pass
Qwen2.5-3B KV cache, 8k ctx, fp16 (36L, 2 KV heads, head 128) 0.3 GB 0.3 GB Pass
Qwen2.5-3B-Instruct FP16 raw forward (all-position logits), RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 6.85GB) 6.8 GB 7.3 GB Pass
Qwen2.5-3B-Instruct FP16 raw forward (all-position logits), RTX 4090 (24GB), 8k ctx, batch 4 (measured peak: 19.773GB; fits - the earlier OOM was the harness holding two logits tensors) 19.8 GB 19.7 GB Pass
DeepSeek-V2-Lite (16B MoE total / 2.4B active) FP16 weights loading on RTX 4090 (24GB) -> OOM Insufficient Insufficient - Pass
Qwen2-VL-2B (2.21B total with auxiliary visual encoder) BF16 empirical weights on RTX 4090 4.1 GB 4.1 GB Pass
Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 15.41GB) 15.4 GB 15.9 GB Pass
Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 4 (measured peak: 18.59GB) 18.6 GB 18.7 GB Pass
Qwen2.5-7B-Instruct FP16 inference, RTX 4090 (24GB), 2k ctx, batch 8 (measured peak: 23.05GB) 23.1 GB 22.8 GB Pass
Qwen2.5-7B-Instruct FP16 empirical model weights on RTX 4090 14.2 GB 14.2 GB Pass
Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 1k ctx, batch 1 (measured peak: 18.40GB) 18.4 GB 19.3 GB Pass
Gemma-2-9B-It FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 19.87GB) 19.9 GB 19.7 GB Pass
Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 1 (measured peak: 17.19GB) 17.2 GB 17.2 GB Pass
Llama-3.1-8B-Instruct FP16 inference, H100 (80GB), 4k ctx, batch 8 (measured peak: 32.73GB) 32.7 GB 30.1 GB Pass
Llama-3.1-8B-Instruct QLoRA r=16 training, 1x RTX 4090 (24GB), bs=2 seq=2048 (measured peak: 19.71GB) 19.7 GB 19.7 GB Pass
Qwen2.5-3B-Instruct full FT, 1x A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 33.59GB) 33.6 GB 32.1 GB Pass
DeepSeek-R1-Distill-Qwen-14B FP16 empirical model weights, A100-SXM4 (80GB) 27.5 GB 27.5 GB Pass
DeepSeek-R1-Distill-Qwen-14B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 30.30GB) 30.3 GB 30.5 GB Pass
DeepSeek-R1-Distill-Qwen-7B FP16 empirical model weights, RTX 4090 (24GB) 14.2 GB 14.2 GB Pass
DeepSeek-R1-Distill-Qwen-7B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 16.23GB) 16.2 GB 15.9 GB Pass
DeepSeek-R1-Distill-Qwen-32B FP16 empirical model weights, A100-SXM4 (80GB) 61.0 GB 61.0 GB Pass
DeepSeek-R1-Distill-Qwen-32B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 66.65GB) 66.7 GB 66.7 GB Pass
Qwen2.5-3B-Instruct FP16 inference, H200 (141GB), 2k ctx, batch 1 (measured peak: 6.76GB) 6.8 GB 6.8 GB Pass
Qwen3.8-27B FP16 empirical model weights, A100-SXM4 (80GB) 50.1 GB 50.1 GB Pass
Qwen3.8-27B FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 54.67GB) 54.7 GB 54.7 GB Pass
Ministral-3-8B-Instruct-2512 FP16 empirical model weights, A100-SXM4 (80GB) 16.6 GB 16.6 GB Pass
Ministral-3-8B-Instruct-2512 FP16 inference, A100-SXM4 (80GB), 2k ctx, batch 1 (measured peak: 20.38GB) 20.4 GB 18.7 GB Pass
Granite-4.2-8B FP16 empirical model weights, RTX 4090 (24GB) 16.4 GB 16.4 GB Pass
Granite-4.2-8B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 18.89GB) 18.9 GB 18.5 GB Pass
Gemma-4-E2B FP16 empirical model weights, RTX 4090 (24GB) 9.5 GB 9.5 GB Pass
Gemma-4-E2B FP16 inference, RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 12.81GB) 12.8 GB 12.4 GB Pass
Gemma-4-12B FP16 empirical model weights, L40S (48GB) 22.3 GB 22.3 GB Pass
Gemma-4-12B FP16 inference, L40S (48GB), 2k ctx, batch 1 (measured peak: 26.62GB) 26.6 GB 25.9 GB Pass
Ministral-3-3B Full FT 32-bit AdamW, A100-SXM4 (80GB), bs=2 seq=2048 (measured peak: 36.95GB) 37.0 GB 70.1 GB Miss
DeepSeek-R1-Distill-Qwen-7B LoRA 32-bit AdamW training, RTX 4090 (24GB) -> OOM Insufficient Insufficient - Pass
Ministral-3-8B-Instruct-2512 LoRA 8-bit AdamW training, RTX 4090 (24GB) -> OOM Insufficient Insufficient - Pass
Qwen3.8-Flash-Next (176.9B) FP16 inference, 2x H200 (141GB) -> OOM (weights exceed cluster capacity) Insufficient Insufficient - Pass
Qwen2.5-3B-Instruct FP16 prefill forward with last-token logits (as generate() does), RTX 4090 (24GB), 2k ctx, batch 1 (measured peak: 6.27GB) 6.3 GB 6.8 GB Pass
Qwen2.5-3B-Instruct FP16 prefill forward with last-token logits (as generate() does), RTX 4090 (24GB), 8k ctx, batch 4 (measured peak: 10.5GB) 10.5 GB 10.4 GB Pass
Gemma-4-12B BF16 inference, 2x B300 SXM6 (288GB), 2k ctx, batch 1 (measured physical peak: 28.72GB) 28.7 GB 26.3 GB Pass
Devstral-2-123B-Instruct FP8 inference, 2x B300 SXM6 (288GB), 2k ctx, batch 1 (measured physical peak: 128.15GB) 128.2 GB 139.8 GB Pass
Qwen2.5-7B-Instruct FP16 eager inference, AMD Instinct MI350X (288GB), 2k ctx, batch 1 (measured physical peak: 16.03GB) 16.0 GB 15.9 GB Pass
Qwen2.5-32B-Instruct BF16, vLLM 0.30.0 defaults, 2x H100 SXM (80GB) TP=2: KV cache blocks after profiling (19691 x 16 tokens, gpu_memory_utilization 0.92) 19691.0 GB 19847.0 GB Pass
Qwen2.5-3B-Instruct BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (25334 x 16 tokens, gpu_memory_utilization 0.92) 25334.0 GB 25158.0 GB Pass
Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (5422 x 16 tokens, gpu_memory_utilization 0.92) 5422.0 GB 6225.0 GB Pass
DeepSeek-R1-Distill-Qwen-7B BF16, vLLM 0.30.0 defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (6567 x 16 tokens, gpu_memory_utilization 0.92) 6567.0 GB 6225.0 GB Pass
Qwen2.5-14B-Instruct BF16, vLLM 0.30.0 defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (14195 x 16 tokens, gpu_memory_utilization 0.92) 14195.0 GB 14824.0 GB Pass
Qwen2.5-14B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (13630 x 16 tokens, gpu_memory_utilization 0.92) 13630.0 GB 14704.0 GB Pass
Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (64060 x 16 tokens, gpu_memory_utilization 0.92) 64060.0 GB 66226.0 GB Pass
Llama-3.1-8B-Instruct BF16, vLLM 0.30.0 defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (27701 x 16 tokens, gpu_memory_utilization 0.92) 27701.0 GB 28554.0 GB Pass
Llama-3.1-8B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 14.07 GiB 14.1 GB 15.4 GB Pass
Llama-3.1-8B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOM Insufficient Insufficient - Pass
Llama-3.1-8B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 15.06 GiB 15.1 GB 16.9 GB Miss
Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 13.87 GiB 13.9 GB 14.2 GB Pass
Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 21.11 GiB 21.1 GB 19.6 GB Pass
Qwen2.5-3B-Instruct LORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOM Insufficient Insufficient - Pass
Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 16.45 GiB 16.4 GB 16.4 GB Pass
Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 22.82 GiB 22.8 GB 23.9 GB Pass
Qwen2.5-3B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs4 seq512: peak device memory 21.11 GiB 21.1 GB 19.6 GB Pass
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 19.91 GiB 19.9 GB 21.6 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq1024: peak device memory 14.43 GiB 14.4 GB 15.1 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 22.69 GiB 22.7 GB 21.9 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs2 seq2048 -> OOM Insufficient Insufficient - Pass
Qwen2.5-7B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 15.7 GiB 15.7 GB 17.1 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, hf_trainer None defaults, 1x RTX 4090 (24GB), bs4 seq2048 -> OOM Insufficient Insufficient - Pass
Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 10.32 GiB 10.3 GB 10.7 GB Pass
Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 12.88 GiB 12.9 GB 12.9 GB Pass
Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 18.01 GiB 18.0 GB 17.8 GB Pass
Qwen2.5-3B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 18.01 GiB 18.0 GB 17.3 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 18.87 GiB 18.9 GB 19.5 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 21.83 GiB 21.8 GB 22.1 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune None defaults, 1x RTX 4090 (24GB), bs2 seq4096 -> OOM Insufficient Insufficient - Pass
Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 16.46 GiB 16.5 GB 17.0 GB Pass
Llama-3.1-8B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 8.52 GiB 8.5 GB 8.6 GB Pass
Llama-3.1-8B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.55 GiB 7.5 GB 7.7 GB Pass
Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 7.08 GiB 7.1 GB 7.3 GB Pass
Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 7.94 GiB 7.9 GB 8.3 GB Pass
Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.55 GiB 7.5 GB 7.5 GB Pass
Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 7.94 GiB 7.9 GB 8.0 GB Pass
Qwen2.5-3B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 7.94 GiB 7.9 GB 7.9 GB Pass
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 16.26 GiB 16.3 GB 16.4 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 6.89 GiB 6.9 GB 7.2 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq8192: peak device memory 8.01 GiB 8.0 GB 8.2 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 7.38 GiB 7.4 GB 7.4 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq4096: peak device memory 8.01 GiB 8.0 GB 8.0 GB Pass
Qwen2.5-7B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 8.01 GiB 8.0 GB 7.8 GB Pass
Phi-3-mini-4k-instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs1 seq2048: peak device memory 17.45 GiB 17.4 GB 18.4 GB Pass
Phi-3-mini-4k-instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 10.91 GiB 10.9 GB 13.8 GB Miss
Phi-3-mini-4k-instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x RTX 4090 (24GB), bs2 seq1024: peak device memory 12.38 GiB 12.4 GB 13.3 GB Pass
Phi-3-mini-4k-instruct LORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 9.79 GiB 9.8 GB 12.7 GB Miss
Phi-3-mini-4k-instruct QLORA r=16, torchtune 0.6.1 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 5.13 GiB 5.1 GB 8.2 GB Miss
Phi-3-mini-4k-instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs1 seq4096: peak device memory 8.64 GiB 8.6 GB 9.0 GB Pass
Phi-3-mini-4k-instruct LORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs2 seq2048: peak device memory 8.64 GiB 8.6 GB 8.9 GB Pass
Phi-3-mini-4k-instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x RTX 4090 (24GB), bs4 seq2048: peak device memory 4.77 GiB 4.8 GB 5.5 GB Miss
Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 50.11 GiB 50.1 GB 45.5 GB Pass
Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 50.7 GiB 50.7 GB 50.9 GB Pass
Qwen2.5-14B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs4 seq2048: peak device memory 52.25 GiB 52.3 GB 52.2 GB Pass
Qwen2.5-14B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 34.37 GiB 34.4 GB 34.4 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq1024: peak device memory 41.59 GiB 41.6 GB 40.9 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 31.53 GiB 31.5 GB 34.5 GB Pass
Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 31.62 GiB 31.6 GB 36.8 GB Miss
Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 44.15 GiB 44.1 GB 47.5 GB Pass
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 32.39 GiB 32.4 GB 30.9 GB Pass
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 49.67 GiB 49.7 GB 44.4 GB Miss
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x A100-SXM4 (80GB), bs4 seq1024: peak device memory 49.67 GiB 49.7 GB 44.4 GB Miss
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 53.9 GiB 53.9 GB 50.5 GB Pass
Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 37.88 GiB 37.9 GB 37.7 GB Pass
Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 59.43 GiB 59.4 GB 60.8 GB Pass
Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 36.38 GiB 36.4 GB 35.3 GB Pass
Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 49.55 GiB 49.5 GB 49.5 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 21.88 GiB 21.9 GB 22.1 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 39.11 GiB 39.1 GB 38.4 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 39.11 GiB 39.1 GB 37.4 GB Pass
Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 19.74 GiB 19.7 GB 19.0 GB Pass
Qwen2.5-14B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 29.98 GiB 30.0 GB 30.6 GB Pass
Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq8192: peak device memory 15.3 GiB 15.3 GB 14.9 GB Pass
Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq2048: peak device memory 13.02 GiB 13.0 GB 12.5 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs1 seq2048: peak device memory 21.08 GiB 21.1 GB 21.4 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq4096: peak device memory 22.75 GiB 22.8 GB 23.1 GB Pass
Qwen2.5-3B-Instruct full FT, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 20.95 GiB 20.9 GB 33.7 GB Miss
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs1 seq16384: peak device memory 18.95 GiB 18.9 GB 19.5 GB Pass
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs2 seq2048: peak device memory 16.3 GiB 16.3 GB 16.4 GB Pass
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs4 seq4096: peak device memory 18.95 GiB 18.9 GB 18.0 GB Pass
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x A100-SXM4 (80GB), bs8 seq2048: peak device memory 18.95 GiB 18.9 GB 17.8 GB Pass
Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 52.31 GiB 52.3 GB 45.5 GB Miss
Llama-3.1-8B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 51.45 GiB 51.5 GB 50.9 GB Pass
Qwen2.5-14B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs4 seq2048: peak device memory 52.06 GiB 52.1 GB 52.2 GB Pass
Qwen2.5-14B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs1 seq2048: peak device memory 34.64 GiB 34.6 GB 34.4 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs1 seq1024: peak device memory 41.81 GiB 41.8 GB 40.9 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 31.76 GiB 31.8 GB 34.5 GB Pass
Qwen2.5-3B-Instruct full FT, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 44.41 GiB 44.4 GB 47.5 GB Pass
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 49.93 GiB 49.9 GB 44.4 GB Miss
Qwen2.5-7B-Instruct LORA r=16, hf_trainer 5.17.0 defaults + grad ckpt, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 54.13 GiB 54.1 GB 50.5 GB Pass
Qwen2.5-14B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 59.28 GiB 59.3 GB 60.8 GB Pass
Qwen2.5-3B-Instruct full FT, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 36.55 GiB 36.5 GB 35.3 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 22.06 GiB 22.1 GB 22.1 GB Pass
Qwen2.5-7B-Instruct LORA r=16, torchtune 0.6.1 defaults, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 39.08 GiB 39.1 GB 37.4 GB Pass
Llama-3.1-8B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 20.03 GiB 20.0 GB 19.0 GB Pass
Qwen2.5-14B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq2048: peak device memory 13.33 GiB 13.3 GB 12.5 GB Pass
Qwen2.5-32B-Instruct QLORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq4096: peak device memory 23.14 GiB 23.1 GB 23.1 GB Pass
Qwen2.5-3B-Instruct full FT, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 21.61 GiB 21.6 GB 33.7 GB Miss
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs2 seq2048: peak device memory 16.6 GiB 16.6 GB 16.4 GB Pass
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs4 seq4096: peak device memory 19.25 GiB 19.3 GB 18.0 GB Pass
Qwen2.5-7B-Instruct LORA r=16, unsloth 2026.9.11 defaults, 1x H100 SXM (80GB), bs8 seq2048: peak device memory 19.25 GiB 19.3 GB 17.8 GB Pass
Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 8192, 1x RTX 4090 (24GB): device memory 5.86 GiB (model 4403 + KV 1024 + compute 116 MiB) 5.9 GB 6.0 GB Pass
Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 8.88 GiB (model 4403 + KV 4096 + compute 140 MiB) 8.9 GB 9.0 GB Pass
Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 65536, 1x RTX 4090 (24GB): device memory 12.92 GiB (model 4403 + KV 8192 + compute 172 MiB) 12.9 GB 13.0 GB Pass
Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.737 I srv llama_server: initializing ... defaults, -c 131072, 1x RTX 4090 (24GB): device memory 20.98 GiB (model 4403 + KV 16384 + compute 236 MiB) 21.0 GB 21.0 GB Pass
Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 9.27 GiB (model 8148 + KV 768 + compute 115 MiB) 9.3 GB 9.3 GB Pass
Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 16384, 1x RTX 4090 (24GB): device memory 11.53 GiB (model 8148 + KV 3072 + compute 127 MiB) 11.5 GB 11.5 GB Pass
Qwen2.5-14B-Instruct Q4_K_M, llama_cpp 0.00.000.811 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 14.55 GiB (model 8148 + KV 6144 + compute 143 MiB) 14.6 GB 14.5 GB Pass
Qwen2.5-3B-Instruct Q8_0, llama_cpp 0.00.000.786 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 3.73 GiB (model 3128 + KV 144 + compute 81 MiB) 3.7 GB 3.7 GB Pass
Qwen2.5-3B-Instruct Q8_0, llama_cpp 0.00.000.786 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 4.74 GiB (model 3128 + KV 1152 + compute 109 MiB) 4.7 GB 4.7 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 4.88 GiB (model 4168 + KV 224 + compute 136 MiB) 4.9 GB 5.0 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 16384, 1x RTX 4090 (24GB): device memory 5.54 GiB (model 4168 + KV 896 + compute 148 MiB) 5.5 GB 5.7 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.862 I srv llama_server: initializing ... defaults, -c 32768, 1x RTX 4090 (24GB): device memory 6.44 GiB (model 4168 + KV 1792 + compute 164 MiB) 6.4 GB 6.5 GB Pass
Llama-3.1-8B-Instruct (llama3.1:8b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x RTX 4090 (24GB): device memory 9.02 GiB 9.0 GB 9.1 GB Pass
Llama-3.1-8B-Instruct (llama3.1:8b), ollama ollama version is 0.34.4 defaults, num_ctx 8192, 1x RTX 4090 (24GB): device memory 5.98 GiB 6.0 GB 6.1 GB Pass
Qwen2.5-14B-Instruct (qwen2.5:14b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 9.27 GiB 9.3 GB 9.4 GB Pass
Qwen2.5-14B-Instruct (qwen2.5:14b), ollama ollama version is 0.34.4 defaults, num_ctx 16384, 1x RTX 4090 (24GB): device memory 11.67 GiB 11.7 GB 11.7 GB Pass
Qwen2.5-3B-Instruct (qwen2.5:3b-instruct-q8_0), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 3.73 GiB 3.7 GB 3.7 GB Pass
Qwen2.5-3B-Instruct (qwen2.5:3b-instruct-q8_0), ollama ollama version is 0.34.4 defaults, num_ctx 32768, 1x RTX 4090 (24GB): device memory 4.85 GiB 4.8 GB 4.7 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x RTX 4090 (24GB): device memory 6.61 GiB 6.6 GB 6.6 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x RTX 4090 (24GB): device memory 4.88 GiB 4.9 GB 5.1 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 16384, 1x RTX 4090 (24GB): device memory 5.7 GiB 5.7 GB 5.7 GB Pass
Phi-3-mini-4k-instruct Q4_K_M, llama_cpp 0.00.001.232 I srv llama_server: initializing ... defaults, -c 2048, 1x RTX 4090 (24GB): device memory 3.44 GiB (model 2229 + KV 768 + compute 62 MiB) 3.4 GB 3.7 GB Pass
Phi-3-mini-4k-instruct Q4_K_M, llama_cpp 0.00.001.232 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 4.19 GiB (model 2229 + KV 1536 + compute 64 MiB) 4.2 GB 4.5 GB Pass
Phi-3-mini-4k-instruct Q8_0, llama_cpp 0.00.000.784 I srv llama_server: initializing ... defaults, -c 4096, 1x RTX 4090 (24GB): device memory 5.7 GiB (model 3773 + KV 1536 + compute 64 MiB) 5.7 GB 5.9 GB Pass
Phi-3-mini-4k-instruct (phi3:3.8b-mini-4k-instruct-q4_K_M), ollama ollama version is 0.34.4 defaults, default context (4096 for this VRAM tier), 1x RTX 4090 (24GB): device memory 4.19 GiB 4.2 GB 4.5 GB Pass
Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.965 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 8.93 GiB (model 4403 + KV 4096 + compute 140 MiB) 8.9 GB 9.0 GB Pass
Llama-3.1-8B-Instruct Q4_K_M, llama_cpp 0.00.000.965 I srv llama_server: initializing ... defaults, -c 131072, 1x A100-SXM4 (80GB): device memory 21.02 GiB (model 4403 + KV 16384 + compute 236 MiB) 21.0 GB 21.0 GB Pass
Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.667 I srv llama_server: initializing ... defaults, -c 4096, 1x A100-SXM4 (80GB): device memory 19.77 GiB (model 18508 + KV 1024 + compute 196 MiB) 19.8 GB 19.3 GB Pass
Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.667 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 26.79 GiB (model 18508 + KV 8192 + compute 224 MiB) 26.8 GB 26.3 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.002.264 I srv llama_server: initializing ... defaults, -c 4096, 1x A100-SXM4 (80GB): device memory 4.91 GiB (model 4168 + KV 224 + compute 136 MiB) 4.9 GB 5.0 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.002.264 I srv llama_server: initializing ... defaults, -c 32768, 1x A100-SXM4 (80GB): device memory 6.88 GiB (model 4168 + KV 1792 + compute 164 MiB) 6.9 GB 6.5 GB Pass
Qwen2.5-32B-Instruct (qwen2.5:32b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x A100-SXM4 (80GB): device memory 27.02 GiB 27.0 GB 26.5 GB Pass
Qwen2.5-32B-Instruct (qwen2.5:32b), ollama ollama version is 0.34.4 defaults, num_ctx 8192, 1x A100-SXM4 (80GB): device memory 20.97 GiB 21.0 GB 20.5 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, default context (32768 for this VRAM tier), 1x A100-SXM4 (80GB): device memory 6.64 GiB 6.6 GB 6.6 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x A100-SXM4 (80GB): device memory 4.89 GiB 4.9 GB 5.1 GB Pass
Qwen2.5-32B-Instruct Q4_K_M, llama_cpp 0.00.000.557 I srv llama_server: initializing ... defaults, -c 8192, 1x H100 SXM (80GB): device memory 20.88 GiB (model 18508 + KV 2048 + compute 200 MiB) 20.9 GB 20.3 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.683 I srv llama_server: initializing ... defaults, -c 4096, 1x H100 SXM (80GB): device memory 5.03 GiB (model 4168 + KV 224 + compute 136 MiB) 5.0 GB 5.0 GB Pass
Qwen2.5-7B-Instruct Q4_K_M, llama_cpp 0.00.000.683 I srv llama_server: initializing ... defaults, -c 32768, 1x H100 SXM (80GB): device memory 6.58 GiB (model 4168 + KV 1792 + compute 164 MiB) 6.6 GB 6.5 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 4096, 1x H100 SXM (80GB): device memory 5.03 GiB 5.0 GB 5.1 GB Pass
Qwen2.5-7B-Instruct (qwen2.5:7b), ollama ollama version is 0.34.4 defaults, num_ctx 32768, 1x H100 SXM (80GB): device memory 6.75 GiB 6.8 GB 6.6 GB Pass
Qwen2.5-3B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (12733 x 32 tokens, free_gpu_memory_fraction 0.9) 12733.0 GB 12296.0 GB Pass
Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (3458 x 32 tokens, free_gpu_memory_fraction 0.9) 3458.0 GB 3465.0 GB Pass
Phi-3-mini-4k-instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x RTX 4090 (24GB): KV cache blocks after profiling (1116 x 32 tokens, free_gpu_memory_fraction 0.9) 1116.0 GB 1056.0 GB Pass
Qwen2.5-14B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (7561 x 32 tokens, free_gpu_memory_fraction 0.9) 7561.0 GB 7550.0 GB Pass
Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (32853 x 32 tokens, free_gpu_memory_fraction 0.9) 32853.0 GB 32872.0 GB Pass
Llama-3.1-8B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (14132 x 32 tokens, free_gpu_memory_fraction 0.9) 14132.0 GB 14192.0 GB Pass
Qwen2.5-14B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (7475 x 32 tokens, free_gpu_memory_fraction 0.9) 7475.0 GB 7550.0 GB Pass
Qwen2.5-7B-Instruct BF16, TensorRT-LLM 1.2.1 LLM API defaults, 1x H100 SXM (80GB): KV cache blocks after profiling (32498 x 32 tokens, free_gpu_memory_fraction 0.9) 32498.0 GB 32872.0 GB Pass
Qwen2.5-7B-Instruct BF16, vLLM 0.30.0 defaults, 1x A100-SXM4 (80GB): KV cache blocks after profiling (65412 x 16 tokens, gpu_memory_utilization 0.92) 65412.0 GB 66519.0 GB Pass