趋近智
The Apple M5 Max (64GB) uses unified memory architecture. For local LLM inference, macOS allocates up to 75% (48 GB) of total system RAM for GPU model weights and KV cache, with 614 GB/s unified memory bandwidth.
总显存
64 GB
可用显存上限
48 GB (75% max allocated to GPU)
显存带宽
614 GB/s
最大推理模型级别
32B Models (Q4)
热设计功耗
150W
随上下文增长模拟的显存需求。当曲线超过参考线时将导致显存不足(OOM)。
量化精度:
不同量化精度下可运行的最大模型参数级别。
上下文长度:
4位量化
32B Models
8位精度
32B Models
16位全精度
14B Models
本地模型训练与适配器微调支持能力。
微调上下文:
QLoRA
32B Models
LoRA
14B Models
全参数微调
7B-8B Models
Guidance for optimal precision and operational ceilings on Apple M5 Max (64GB).
Inference Sweet Spot
Optimized for 14B models at uncompressed or Q8 precision, and 32B models at Q4 with up to 32k context.
Bandwidth & Speed Profile
With 614 GB/s aggregate bandwidth, batch size 1 inference operates in a memory-bandwidth bound regime. Generates approximately 136 tok/s on an 8B Q4 model and 16 tok/s on a 70B Q4 model.
Unified Memory Headroom
macOS limits GPU memory allocation to approximately 75% of physical unified RAM by default to ensure system stability.
APX AI
在线