The Apple M3 Ultra (256GB) uses a unified memory architecture. For reliable local LLM inference without OS memory pressure, this benchmark models a 75% usable allocation (192 GB) with 819 GB/s unified memory bandwidth.
Memory
256 GB
Usable Memory Ceiling
192 GB (75% usable budget)
Bandwidth
819.2 GB/s
Max Inference Class
70B+ Models (Q4)
TDP
250W
Estimated memory requirements as context length increases. Curves crossing above the reference line indicate Out of Memory (OOM).
Precision:
Inference throughput, memory compatibility, and generation speeds for open-weights LLMs on Apple M3 Ultra (256GB).
Sort by:
Context:
KV Cache:
| Rank | Model | Parameters | Precision | TPS (Approx.) | TTFT | Can Run |
|---|---|---|---|---|---|---|
#30 | DeepSeek-V4-Flash-Vision-Exp | 290.9B (20.4B active) | INT4 | 52 tok/s | ~6.9s | |
#33 | Qwen 3.8 27B | 27B | INT4 Q8 FP16 | 39 tok/s 23 tok/s 12 tok/s | ~9.5s ~9.6s ~9.6s | |
#40 | MiMo-V2.6-Flash | 309B | INT4 | 4 tok/s | ~116.3s | |
#41 | Qwen 3.8 Flash | 176B (6B active) | INT4 Q8 | 123 tok/s 79 tok/s | ~2.1s ~2.2s | |
#42 | DeepSeek-V4-Flash | 284B (13B active) | INT4 | 74 tok/s | ~4.4s | |
#43 | K2 Horizon 375B A23B | 375B (23B active) | INT4 | 31 tok/s | ~13.1s | |
#48 | Agnes 3.0 Flash | 33B | INT4 Q8 FP16 | 33 tok/s 19 tok/s 10 tok/s | ~11.6s ~11.6s ~11.7s | |
#61 | Sarvam-105B | 106B (10.3B active) | INT4 Q8 | 83 tok/s 53 tok/s | ~4.9s ~4.9s | |
#72 | K2 Horizon MoVA 36B A4B | 36B (4B active) | INT4 Q8 FP16 | 131 tok/s 99 tok/s 63 tok/s | ~1.9s ~1.9s ~1.9s | |
#75 | Ling 3.0 Flash VL | 124B (5.5B active) | INT4 Q8 | 120 tok/s 84 tok/s | ~3.5s ~3.5s | |
#80 | GLM-5.3 Flash | 313.3B (17.3B active) | INT4 | 56 tok/s | ~6.4s | |
#82 | Qwen3 235B A22B Thinking | 235B (22B active) | INT4 | 16 tok/s | ~9.6s | |
#83 | Qwen3.5-27B | 27B | INT4 Q8 FP16 | 39 tok/s 23 tok/s 12 tok/s | ~9.5s ~9.6s ~9.6s | |
#84 | Qwen3.5-122B-A10B | 122B (10B active) | INT4 Q8 | 67 tok/s 47 tok/s | ~5.0s ~5.0s | |
#92 | Sarvam-30B | 32B (2.4B active) | INT4 Q8 FP16 | 159 tok/s 128 tok/s 89 tok/s | ~1.3s ~1.3s ~1.3s | |
#94 | Qwen3.6 35B A3B | 35B (3B active) | INT4 Q8 FP16 | 147 tok/s 115 tok/s 77 tok/s | ~1.5s ~1.5s ~1.5s | |
#96 | Muse Glimmer 30B | 30B | INT4 Q8 FP16 | 36 tok/s 21 tok/s 11 tok/s | ~10.6s ~10.6s ~10.6s | |
#100 | Phi-4 Reasoning Plus | 14B | INT4 Q8 FP16 | 51 tok/s 35 tok/s 21 tok/s | ~5.0s ~5.0s ~5.1s | |
#101 | MiMo V2 Flash | 15B (309B active) | INT4 Q8 FP16 | 4 tok/s 2 tok/s 1 tok/s | ~104.4s ~104.6s ~105.0s | |
#105 | Inkling-Small | 276B (12B active) | INT4 | 15 tok/s | ~7.3s | |
#106 | GLM-4.5-Air | 106B (12B active) | INT4 Q8 | 26 tok/s 22 tok/s | ~5.5s ~5.5s | |
#108 | Gemma 4 26B A4B | 25.2B (3.8B active) | INT4 Q8 FP16 | 138 tok/s 104 tok/s 66 tok/s | ~1.7s ~1.7s ~1.7s | |
#109 | Gemma 4 12B | 11.95B | INT4 Q8 FP16 | 57 tok/s 40 tok/s 24 tok/s | ~4.3s ~4.3s ~4.3s | |
#111 | MiniMax M2 | 229B (10B active) | INT4 | 20 tok/s | ~5.5s | |
#113 | Qwen3.5-9B | 9B | INT4 Q8 FP16 | 92 tok/s 60 tok/s 34 tok/s | ~3.3s ~3.3s ~3.3s | |
#114 | Qwen3.5-35B-A3B | 35B (3B active) | INT4 Q8 FP16 | 147 tok/s 115 tok/s 77 tok/s | ~1.5s ~1.5s ~1.5s | |
#114 | Ling 3.0 Flash | 124B (5.1B active) | INT4 Q8 | 125 tok/s 89 tok/s | ~3.3s ~3.4s | |
#118 | K2 Horizon 7B | 7B | INT4 Q8 FP16 | 109 tok/s 73 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
#119 | Step 3.5 Flash | 196.81B (11B active) | INT4 Q8 | 62 tok/s 39 tok/s | ~5.5s ~6.5s | |
#126 | Qwen3 Next 80B A3B | 80B (79B active) | INT4 Q8 FP16 | 12 tok/s 7 tok/s 4 tok/s | ~27.7s ~27.7s ~30.2s | |
#128 | NVIDIA Nemotron 3 Nano 30B-A3B | 3.5B (30B active) | INT4 Q8 FP16 | 34 tok/s 21 tok/s 11 tok/s | ~10.2s ~10.3s ~10.3s | |
#129 | Gemma 4 31B | 30.7B | INT4 Q8 FP16 | 35 tok/s 21 tok/s 11 tok/s | ~10.8s ~10.8s ~10.9s | |
#130 | GPT-OSS 120B | 117B (5.1B active) | INT4 Q8 | 27 tok/s 27 tok/s | ~3.3s ~3.3s | |
#131 | Nemotron 3.5 Lightning | 30B (3B active) | INT4 Q8 FP16 | 69 tok/s 64 tok/s 50 tok/s | ~1.5s ~1.5s ~1.5s | |
#132 | Qwen3.5-4B | 4B | INT4 Q8 FP16 | 148 tok/s 107 tok/s 66 tok/s | ~1.5s ~1.5s ~1.5s | |
#133 | Qwen3-30B-A3B | 30B (3B active) | INT4 Q8 FP16 | 148 tok/s 116 tok/s 77 tok/s | ~1.5s ~1.5s ~1.5s | |
#136 | GPT-OSS 20B | 21B (3.6B active) | INT4 Q8 FP16 | 76 tok/s 67 tok/s 50 tok/s | ~1.6s ~1.6s ~1.6s | |
#139 | Qwen3-235B-A22B | 235B (22B active) | INT4 | 39 tok/s | ~9.6s | |
#144 | Mistral Medium 3.5 128B | 128B | INT4 Q8 | 7 tok/s 5 tok/s | ~44.8s ~44.8s | |
#146 | Granite 4.2 8B | 8B | INT4 Q8 FP16 | 100 tok/s 66 tok/s 37 tok/s | ~2.9s ~2.9s ~2.9s | |
#150 | Gemma 4 E4B | 8B | INT4 Q8 FP16 | 100 tok/s 66 tok/s 37 tok/s | ~2.9s ~2.9s ~2.9s | |
#151 | Nemotron 3.5 Lightning 30B A3B | 30B (3B active) | INT4 Q8 FP16 | 148 tok/s 116 tok/s 77 tok/s | ~1.5s ~1.5s ~1.5s | |
#153 | Qwen3-32B | 32B | INT4 Q8 FP16 | 34 tok/s 20 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#154 | Qwen2.5-72B | 72B | INT4 Q8 FP16 | 16 tok/s 9 tok/s 5 tok/s | ~25.2s ~25.2s ~25.3s | |
#155 | Mistral Large 3 | 41B (675B active) | INT4 Q8 FP16 | 2 tok/s 1 tok/s 1 tok/s | ~227.8s ~228.3s ~229.2s | |
#156 | GLM-4 | 32B | INT4 Q8 FP16 | 27 tok/s 17 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#157 | DeepSeek-R1 70B | 70B | INT4 Q8 FP16 | 17 tok/s 10 tok/s 5 tok/s | ~24.5s ~24.5s ~24.6s | |
#158 | GLM-4V | 9B | INT4 Q8 FP16 | 70 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
#159 | Gemma 3 27B | 27B | INT4 Q8 FP16 | 39 tok/s 23 tok/s 12 tok/s | ~9.5s ~9.6s ~9.6s | |
#160 | Phi-4 | 14B | INT4 Q8 FP16 | 67 tok/s 42 tok/s 23 tok/s | ~5.0s ~5.0s ~5.0s | |
#161 | Llama 3.3 70B | 70B | INT4 Q8 FP16 | 17 tok/s 9 tok/s 5 tok/s | ~24.5s ~24.5s ~24.6s | |
#162 | DeepSeek-R1 32B | 32B | INT4 Q8 FP16 | 35 tok/s 20 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#163 | Hunyuan Standard | 52B (389B active) | INT4 Q8 FP16 | 3 tok/s 2 tok/s 1 tok/s | ~131.6s ~131.9s ~132.4s | |
#164 | Qwen2.5-32B | 32B | INT4 Q8 FP16 | 34 tok/s 20 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#168 | Magistral Small | 24B | INT4 Q8 FP16 | 34 tok/s 22 tok/s 13 tok/s | ~8.5s ~8.5s ~8.6s | |
#169 | OLMo 3 7B Think | 7B | INT4 Q8 FP16 | 83 tok/s 61 tok/s 38 tok/s | ~2.5s ~2.6s ~2.6s | |
#170 | OLMo 3.1 32B Instruct | 32B | INT4 Q8 FP16 | 27 tok/s 17 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#171 | Llama 3.1 70B | 70B | INT4 Q8 FP16 | 17 tok/s 9 tok/s 5 tok/s | ~24.5s ~24.5s ~24.6s | |
#172 | DeepSeek-R1 14B | 14B | INT4 Q8 FP16 | 69 tok/s 42 tok/s 23 tok/s | ~5.0s ~5.0s ~5.0s | |
#173 | Ministral 3 14B | 14B | INT4 Q8 FP16 | 51 tok/s 35 tok/s 21 tok/s | ~5.0s ~5.0s ~5.1s | |
#174 | Mistral-Large-2407 | 123B | INT4 Q8 | 10 tok/s 5 tok/s | ~43.0s ~43.1s | |
#175 | Llama 4 Scout | 109B (17B active) | INT4 Q8 | 50 tok/s 32 tok/s | ~7.2s ~7.2s | |
#177 | Gemma 3 12B | 12B | INT4 Q8 FP16 | 75 tok/s 47 tok/s 26 tok/s | ~4.3s ~4.3s ~4.3s | |
#178 | Command A | 111B | INT4 Q8 | 9 tok/s 5 tok/s | ~38.9s ~39.0s | |
#179 | Qwen2-72B | 72B | INT4 Q8 FP16 | 16 tok/s 9 tok/s 5 tok/s | ~25.2s ~25.2s ~25.3s | |
#180 | OLMo 3 32B Base | 32B | INT4 Q8 FP16 | 27 tok/s 17 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#181 | Mistral-Small-2501 | 24B | INT4 Q8 FP16 | 43 tok/s 26 tok/s 14 tok/s | ~8.5s ~8.5s ~8.6s | |
#182 | Ministral 3 8B | 8B | INT4 Q8 FP16 | 76 tok/s 55 tok/s 34 tok/s | ~2.9s ~2.9s ~2.9s | |
#183 | Qwen3.5-2B | 2B | INT4 Q8 FP16 | 195 tok/s 156 tok/s 108 tok/s | ~794 ms ~796 ms ~799 ms | |
#184 | MiniCPM5-2B | 2B | INT4 Q8 FP16 | 195 tok/s 156 tok/s 108 tok/s | ~794 ms ~796 ms ~799 ms | |
#185 | Qwen2.5-14B | 14B | INT4 Q8 FP16 | 67 tok/s 42 tok/s 23 tok/s | ~5.0s ~5.0s ~5.0s | |
#186 | Gemma 2 27B | 27B | INT4 Q8 FP16 | 39 tok/s 23 tok/s 12 tok/s | ~9.5s ~9.6s ~9.6s | |
#187 | OLMo 3.1 32B Think | 32B | INT4 Q8 FP16 | 27 tok/s 17 tok/s 10 tok/s | ~11.3s ~11.3s ~11.3s | |
#188 | Gemma 4 E2B | 5.1B | INT4 Q8 FP16 | 131 tok/s 91 tok/s 55 tok/s | ~1.9s ~1.9s ~1.9s | |
#189 | Qwen3-4B | 4B | INT4 Q8 FP16 | 148 tok/s 107 tok/s 66 tok/s | ~1.5s ~1.5s ~1.5s | |
#190 | DeepSeek-R1 8B | 8B | INT4 Q8 FP16 | 102 tok/s 67 tok/s 38 tok/s | ~2.9s ~2.9s ~2.9s | |
#191 | Llama 3 70B | 70B | INT4 Q8 FP16 | 17 tok/s 9 tok/s 5 tok/s | ~24.5s ~24.5s ~24.6s | |
#192 | Command R Plus | 104B | INT4 Q8 | 9 tok/s 6 tok/s | ~36.4s ~36.5s | |
#193 | Ministral 3 3B | 3B | INT4 Q8 FP16 | 126 tok/s 103 tok/s 71 tok/s | ~1.1s ~1.1s ~1.2s | |
#194 | ERNIE-4.5-300B-A47B | 300B (47B active) | INT4 | 20 tok/s | ~20.3s | |
#195 | Qwen3-14B | 14B | INT4 Q8 FP16 | 67 tok/s 42 tok/s 23 tok/s | ~5.0s ~5.0s ~5.0s | |
#196 | Kimi Linear 48B A3B Instruct | 48B (3B active) | INT4 Q8 FP16 | 60 tok/s 56 tok/s 45 tok/s | ~1.6s ~1.6s ~1.6s | |
#197 | Gemma 2 9B | 9B | INT4 Q8 FP16 | 92 tok/s 60 tok/s 34 tok/s | ~3.3s ~3.3s ~3.3s | |
#198 | Qwen2.5-7B | 7B | INT4 Q8 FP16 | 109 tok/s 73 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
#199 | Qwen3-8B | 8B | INT4 Q8 FP16 | 100 tok/s 66 tok/s 37 tok/s | ~2.9s ~2.9s ~2.9s | |
#200 | Gemma 3 4B | 4B | INT4 Q8 FP16 | 148 tok/s 107 tok/s 66 tok/s | ~1.5s ~1.5s ~1.5s | |
#201 | DeepSeek-R1 7B | 7B | INT4 Q8 FP16 | 111 tok/s 74 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
#202 | Llama 3.1 8B | 8B | INT4 Q8 FP16 | 100 tok/s 66 tok/s 37 tok/s | ~2.9s ~2.9s ~2.9s | |
#203 | Mixtral-8x22B-v0.1 | 176B (22B active) | INT4 Q8 | 41 tok/s 25 tok/s | ~9.0s ~9.8s | |
#204 | Qwen3-1.7B | 1.7B | INT4 Q8 FP16 | 206 tok/s 168 tok/s 119 tok/s | ~680 ms ~681 ms ~684 ms | |
#205 | Command R | 35B | INT4 Q8 FP16 | 25 tok/s 16 tok/s 9 tok/s | ~12.3s ~12.3s ~12.4s | |
#206 | Gemma 3 270M | 0.27B | INT4 Q8 FP16 | 269 tok/s 257 tok/s 234 tok/s | ~111 ms ~111 ms ~111 ms | |
#208 | Llama 3 8B | 8B | INT4 Q8 FP16 | 100 tok/s 66 tok/s 37 tok/s | ~2.9s ~2.9s ~2.9s | |
#209 | Qwen3-0.6B | 0.6B | INT4 Q8 FP16 | 260 tok/s 236 tok/s 196 tok/s | ~242 ms ~242 ms ~243 ms | |
#210 | Phi-4-Mini | 3.8B | INT4 Q8 FP16 | 152 tok/s 111 tok/s 69 tok/s | ~1.4s ~1.4s ~1.4s | |
#211 | Mixtral-8x7B-v0.1 | 46.7B (7B active) | INT4 Q8 FP16 | 99 tok/s 68 tok/s 41 tok/s | ~3.0s ~3.0s ~3.0s | |
#212 | Ministral-8B-2410 | 8B | INT4 Q8 FP16 | 100 tok/s 66 tok/s 37 tok/s | ~2.9s ~2.9s ~2.9s | |
#213 | DeepSeek-R1 1.5B | 1.5B | INT4 Q8 FP16 | 214 tok/s 177 tok/s 128 tok/s | ~602 ms ~603 ms ~605 ms | |
#214 | Llama 3.2 3B | 3B | INT4 Q8 FP16 | 169 tok/s 127 tok/s 82 tok/s | ~1.1s ~1.1s ~1.1s | |
#215 | Qwen3.5-0.8B | 0.8B | INT4 Q8 FP16 | 248 tok/s 220 tok/s 175 tok/s | ~326 ms ~326 ms ~327 ms | |
#216 | Phi-3-mini | 3.8B | INT4 Q8 FP16 | 152 tok/s 111 tok/s 69 tok/s | ~1.4s ~1.4s ~1.4s | |
#217 | Gemma 3n E2B IT | 6B (2B active) | INT4 Q8 FP16 | 126 tok/s 111 tok/s 84 tok/s | ~851 ms ~852 ms ~855 ms | |
#218 | Mistral-7B-Instruct-v0.1 | 7.3B | INT4 Q8 FP16 | 106 tok/s 71 tok/s 41 tok/s | ~2.6s ~2.7s ~2.7s | |
#219 | Llama 3.2 1B | 1B | INT4 Q8 FP16 | 236 tok/s 205 tok/s 158 tok/s | ~413 ms ~413 ms ~415 ms | |
#220 | Gemma 2 2B | 2B | INT4 Q8 FP16 | 195 tok/s 156 tok/s 108 tok/s | ~794 ms ~796 ms ~799 ms | |
#221 | Phi-3-medium | 14B | INT4 Q8 FP16 | 67 tok/s 42 tok/s 23 tok/s | ~5.0s ~5.0s ~5.0s | |
#222 | OLMo 3 7B Instruct | 7B | INT4 Q8 FP16 | 83 tok/s 61 tok/s 38 tok/s | ~2.5s ~2.6s ~2.6s | |
#223 | Yi-34B | 34B | INT4 Q8 FP16 | 26 tok/s 16 tok/s 9 tok/s | ~12.0s ~12.0s ~12.0s | |
#224 | Qwen2-7B | 7B | INT4 Q8 FP16 | 109 tok/s 73 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
#225 | Phi-3-small | 7B | INT4 Q8 FP16 | 109 tok/s 73 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
#226 | Gemma 3 1B | 1B | INT4 Q8 FP16 | 236 tok/s 205 tok/s 158 tok/s | ~413 ms ~413 ms ~415 ms | |
#227 | Mistral-7B-Instruct-v0.2 | 7.3B | INT4 Q8 FP16 | 106 tok/s 71 tok/s 41 tok/s | ~2.6s ~2.7s ~2.7s | |
#228 | Falcon-180B | 180B | INT4 Q8 | 7 tok/s 4 tok/s | ~62.3s ~67.8s | |
#229 | Gemma 1 7B | 7B | INT4 Q8 FP16 | 83 tok/s 61 tok/s 38 tok/s | ~2.5s ~2.6s ~2.6s | |
#230 | Gemma 1 2B | 2B | INT4 Q8 FP16 | 203 tok/s 161 tok/s 110 tok/s | ~794 ms ~796 ms ~798 ms | |
#231 | ChatGLM3-6B | 6B | INT4 Q8 FP16 | 90 tok/s 68 tok/s 43 tok/s | ~2.2s ~2.2s ~2.2s | |
#232 | ChatGLM2-6B | 6B | INT4 Q8 FP16 | 90 tok/s 68 tok/s 43 tok/s | ~2.2s ~2.2s ~2.2s | |
#233 | ChatGLM-6B | 6B | INT4 Q8 FP16 | 90 tok/s 68 tok/s 43 tok/s | ~2.2s ~2.2s ~2.2s | |
Unranked Models (Sorted by Speed) | ||||||
- | Qwen2-0.5B | 0.5B | INT4 Q8 FP16 | 267 tok/s 246 tok/s 209 tok/s | ~201 ms ~202 ms ~202 ms | |
- | Qwen2.5-0.5B | 0.5B | INT4 Q8 FP16 | 267 tok/s 246 tok/s 209 tok/s | ~201 ms ~202 ms ~202 ms | |
- | ERNIE-4.5-0.3B | 0.3B | INT4 Q8 FP16 | 264 tok/s 253 tok/s 228 tok/s | ~123 ms ~123 ms ~124 ms | |
- | ERNIE-4.5-0.3B-Base | 0.3B | INT4 Q8 FP16 | 264 tok/s 253 tok/s 228 tok/s | ~123 ms ~123 ms ~124 ms | |
- | Falcon-1B | 1B | INT4 Q8 FP16 | 243 tok/s 210 tok/s 161 tok/s | ~413 ms ~413 ms ~415 ms | |
- | Falcon3-1B | 1B | INT4 Q8 FP16 | 236 tok/s 205 tok/s 158 tok/s | ~413 ms ~413 ms ~415 ms | |
- | Qwen2-1.5B | 1.5B | INT4 Q8 FP16 | 214 tok/s 177 tok/s 128 tok/s | ~602 ms ~603 ms ~605 ms | |
- | Qwen2.5-1.5B | 1.5B | INT4 Q8 FP16 | 214 tok/s 177 tok/s 128 tok/s | ~602 ms ~603 ms ~605 ms | |
- | Spark X2.5 1.7B | 1.7B | INT4 Q8 FP16 | 206 tok/s 168 tok/s 119 tok/s | ~680 ms ~681 ms ~684 ms | |
- | Hy-MT2 1.8B | 1.8B | INT4 Q8 FP16 | 202 tok/s 164 tok/s 115 tok/s | ~716 ms ~717 ms ~720 ms | |
- | Falcon-3B | 3B | INT4 Q8 FP16 | 175 tok/s 130 tok/s 83 tok/s | ~1.1s ~1.1s ~1.1s | |
- | CroissantLLM Base | 1.3B | INT4 Q8 FP16 | 173 tok/s 154 tok/s 119 tok/s | ~526 ms ~527 ms ~529 ms | |
- | Phi-1 | 1.3B | INT4 Q8 FP16 | 173 tok/s 154 tok/s 119 tok/s | ~526 ms ~527 ms ~529 ms | |
- | Phi-1.5 | 1.3B | INT4 Q8 FP16 | 173 tok/s 154 tok/s 119 tok/s | ~526 ms ~527 ms ~529 ms | |
- | DeepSeek-R1 3B | 3B | INT4 Q8 FP16 | 170 tok/s 128 tok/s 82 tok/s | ~1.1s ~1.1s ~1.1s | |
- | Falcon3-3B | 3B | INT4 Q8 FP16 | 169 tok/s 127 tok/s 82 tok/s | ~1.1s ~1.1s ~1.1s | |
- | Qwen2.5-3B | 3B | INT4 Q8 FP16 | 169 tok/s 127 tok/s 82 tok/s | ~1.1s ~1.1s ~1.1s | |
- | ERNIE-4.5-21B-A3B | 21B (3B active) | INT4 Q8 FP16 | 153 tok/s 119 tok/s 79 tok/s | ~1.4s ~1.4s ~1.4s | |
- | ERNIE-4.5-21B-A3B-Base | 21B (3B active) | INT4 Q8 FP16 | 153 tok/s 119 tok/s 79 tok/s | ~1.4s ~1.4s ~1.4s | |
- | ERNIE-4.5-VL-28B-A3B | 28B (3B active) | INT4 Q8 FP16 | 150 tok/s 117 tok/s 78 tok/s | ~1.5s ~1.5s ~1.5s | |
- | ERNIE-4.5-VL-28B-A3B-Base | 28B (3B active) | INT4 Q8 FP16 | 150 tok/s 117 tok/s 78 tok/s | ~1.5s ~1.5s ~1.5s | |
- | Hy-MT2 30B A3B | 30B (3B active) | INT4 Q8 FP16 | 148 tok/s 116 tok/s 77 tok/s | ~1.5s ~1.5s ~1.5s | |
- | Spark X2.5 4B | 4B | INT4 Q8 FP16 | 148 tok/s 107 tok/s 66 tok/s | ~1.5s ~1.5s ~1.5s | |
- | Phi-2 | 2.7B | INT4 Q8 FP16 | 131 tok/s 108 tok/s 76 tok/s | ~1.0s ~1.0s ~1.0s | |
- | MaLLaM-3B | 3B | INT4 Q8 FP16 | 126 tok/s 103 tok/s 71 tok/s | ~1.1s ~1.1s ~1.2s | |
- | SmolLM3 3B | 3B | INT4 Q8 FP16 | 126 tok/s 103 tok/s 71 tok/s | ~1.1s ~1.1s ~1.2s | |
- | AliceAI Foundation 80B A3B Base | 80B (3B active) | INT4 Q8 FP16 | 124 tok/s 102 tok/s 67 tok/s | ~2.0s ~2.0s ~2.2s | |
- | Falcon-7B | 7B | INT4 Q8 FP16 | 113 tok/s 74 tok/s 43 tok/s | ~2.5s ~2.5s ~2.6s | |
- | Falcon3-7B | 7B | INT4 Q8 FP16 | 109 tok/s 73 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
- | Hy-MT2 7B | 7B | INT4 Q8 FP16 | 109 tok/s 73 tok/s 42 tok/s | ~2.5s ~2.5s ~2.6s | |
- | Mistral-7B-v0.1 | 7.3B | INT4 Q8 FP16 | 106 tok/s 71 tok/s 41 tok/s | ~2.6s ~2.7s ~2.7s | |
- | LensVLM 9B | 9B | INT4 Q8 FP16 | 92 tok/s 60 tok/s 34 tok/s | ~3.3s ~3.3s ~3.3s | |
- | MiMo V2.6 Distill Qwen 9B | 9B | INT4 Q8 FP16 | 92 tok/s 60 tok/s 34 tok/s | ~3.3s ~3.3s ~3.3s | |
- | ChatGLM3-6B-32K | 6B | INT4 Q8 FP16 | 90 tok/s 68 tok/s 43 tok/s | ~2.2s ~2.2s ~2.2s | |
- | Yi-6B | 6B | INT4 Q8 FP16 | 90 tok/s 68 tok/s 43 tok/s | ~2.2s ~2.2s ~2.2s | |
- | Falcon3-10B | 10B | INT4 Q8 FP16 | 86 tok/s 55 tok/s 31 tok/s | ~3.6s ~3.6s ~3.6s | |
- | Kimi-VL-A3B-Instruct | 16B (3B active) | INT4 Q8 FP16 | 86 tok/s 76 tok/s 57 tok/s | ~1.3s ~1.3s ~1.4s | |
- | Falcon2-11B | 11B | INT4 Q8 FP16 | 83 tok/s 52 tok/s 29 tok/s | ~4.0s ~4.0s ~4.0s | |
- | Hunyuan Lite | 7B | INT4 Q8 FP16 | 83 tok/s 61 tok/s 38 tok/s | ~2.5s ~2.6s ~2.6s | |
- | OLMo 3 7B Base | 7B | INT4 Q8 FP16 | 83 tok/s 61 tok/s 38 tok/s | ~2.5s ~2.6s ~2.6s | |
- | SEA-LION-7B | 7.1B | INT4 Q8 FP16 | 82 tok/s 60 tok/s 37 tok/s | ~2.6s ~2.6s ~2.6s | |
- | SEA-LION-7B-Instruct | 7.1B | INT4 Q8 FP16 | 82 tok/s 60 tok/s 37 tok/s | ~2.6s ~2.6s ~2.6s | |
- | Laguna S 2.1 | 118B (8B active) | INT4 Q8 | 76 tok/s 55 tok/s | ~4.3s ~4.3s | |
- | Sahabat-AI-Llama3-8B-Instruct | 8B | INT4 Q8 FP16 | 76 tok/s 55 tok/s 34 tok/s | ~2.9s ~2.9s ~2.9s | |
- | Typhoon-2-8B | 8B | INT4 Q8 FP16 | 76 tok/s 55 tok/s 34 tok/s | ~2.9s ~2.9s ~2.9s | |
- | GLM-4-9B | 9B | INT4 Q8 FP16 | 70 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
- | GLM-4-9B-Chat | 9B | INT4 Q8 FP16 | 70 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
- | GLM-4-9B-Chat-1M | 9B | INT4 Q8 FP16 | 70 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
- | Yi-9B | 9B | INT4 Q8 FP16 | 70 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
- | Sahabat-AI-Gemma2-9B | 9.2B | INT4 Q8 FP16 | 69 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
- | Sahabat-AI-Gemma2-9B-Instruct | 9.2B | INT4 Q8 FP16 | 69 tok/s 50 tok/s 30 tok/s | ~3.3s ~3.3s ~3.3s | |
- | GreenMind-14B-R1 | 14B | INT4 Q8 FP16 | 51 tok/s 35 tok/s 21 tok/s | ~5.0s ~5.0s ~5.1s | |
- | Hunyuan A13B | 80B (13B active) | INT4 Q8 FP16 | 31 tok/s 25 tok/s 16 tok/s | ~5.4s ~5.4s ~5.9s | |
- | Falcon-40B | 40B | INT4 Q8 FP16 | 29 tok/s 16 tok/s 8 tok/s | ~14.0s ~14.0s ~14.1s | |
- | ERNIE-4.5-300B-A47B-Base | 300B (47B active) | INT4 | 20 tok/s | ~20.3s | |
- | Kimi-Dev-72B | 72B | INT4 Q8 FP16 | 13 tok/s 8 tok/s 4 tok/s | ~25.2s ~25.2s ~27.5s | |
- | MiMo V2.6 Flash RL | 309B (15B active) | INT4 | 13 tok/s | ~9.5s | |
- | Typhoon-2-70B | 70B | INT4 Q8 FP16 | 13 tok/s 8 tok/s 5 tok/s | ~24.5s ~24.5s ~24.6s | |
- | GLM-130B | 130B | INT4 Q8 | 7 tok/s 5 tok/s | ~45.4s ~45.5s | |
Maximum model parameter class runnable by weight precision.
Context:
4-bit
70B+ Models
8-bit
70B Models
16-bit
70B Models
Hardware capacity for local training and adapter fine-tuning.
Context Length:
QLoRA
70B+ Models
LoRA
70B Models
Full Parameter Training
7B-8B Models
Guidance for optimal precision and operational ceilings on Apple M3 Ultra (256GB).
Inference Sweet Spot
Suitable for 70B parameter models at Q4/Q8 with extended 32k context windows, or high-throughput batching of smaller dense architectures.
Bandwidth & Speed Profile
With 819.2 GB/s aggregate bandwidth, batch size 1 inference operates in a memory-bandwidth bound regime. Generates approximately 182 tok/s on an 8B Q4 model and 22 tok/s on a 70B Q4 model.
Unified Memory Budget
Modeled with a 75% usable allocation budget to preserve system RAM for macOS and display compositor processes.
©2025 ApX Machine Learning
Assistant
Online