趋近智
The NVIDIA RTX 4060 Ti (16GB) is equipped with 16 GB of dedicated VRAM and 288 GB/s memory bandwidth. It provides 100% addressable VRAM for CUDA and TensorRT-LLM runtimes, supporting full precision and quantized inference.
总显存
16 GB
可用显存上限
16 GB
显存带宽
288 GB/s
最大推理模型级别
14B Models (Q4)
热设计功耗
165W
随上下文增长模拟的显存需求。当曲线超过参考线时将导致显存不足(OOM)。
量化精度:
不同量化精度下可运行的最大模型参数级别。
上下文长度:
4位量化
14B Models
8位精度
8B Models
16位全精度
3B Models
本地模型训练与适配器微调支持能力。
微调上下文:
QLoRA
14B Models
LoRA
8B Models
全参数微调
1B Models
Guidance for optimal precision and operational ceilings on NVIDIA RTX 4060 Ti (16GB).
Inference Sweet Spot
Optimal for 8B models across extended context (32k+). 14B models fit comfortably at Q4 precision with standard context.
Bandwidth & Speed Profile
With 288 GB/s aggregate bandwidth, batch size 1 inference operates in a memory-bandwidth bound regime. Generates approximately 64 tok/s on an 8B Q4 model and 8 tok/s on a 70B Q4 model.
APX AI
在线