ApX logoApX logo

NVIDIA RTX 6000 Ada Generation (48GB)

The NVIDIA RTX 6000 Ada Generation (48GB) is equipped with 48 GB of dedicated VRAM and 960 GB/s memory bandwidth. It provides 100% addressable VRAM for CUDA and TensorRT-LLM runtimes, supporting full precision and quantized inference.

Memory

48 GB

Usable Memory Ceiling

48 GB

Bandwidth

960 GB/s

Max Inference Class

32B Models (Q4)

TDP

300W

VRAM Usage vs Context Length

Simulated memory requirements as context increases. Curves crossing above the reference line will cause Out Of Memory (OOM) errors.

Precision:

Model Compatibility & Speed Matrix

Simulated inference throughput and memory feasibility across popular open-weights models.

Context:

KV Cache:

GPUs:

ModelParametersPrecisionTPS (Approx.)TTFTCan Run
Llama 3.2 1B

1.2B

INT4

191 tok/s

~297 ms

Llama 3.2 1B

1.2B

Q8

143 tok/s

~298 ms

Llama 3.2 1B

1.2B

FP16

92 tok/s

~302 ms

Llama 3.2 3B

3.2B

INT4

110 tok/s

~731 ms

Llama 3.2 3B

3.2B

Q8

73 tok/s

~736 ms

Llama 3.2 3B

3.2B

FP16

41 tok/s

~746 ms

Mistral 7B

7.2B

INT4

61 tok/s

~1573 ms

Mistral 7B

7.2B

Q8

37 tok/s

~1584 ms

Mistral 7B

7.2B

FP16

-

-

Llama 3.1 8B

8.0B

INT4

56 tok/s

~1742 ms

Llama 3.1 8B

8.0B

Q8

32 tok/s

~1905 ms

Llama 3.1 8B

8.0B

FP16

-

-

Mistral Nemo 12B

12.2B

INT4

39 tok/s

~2636 ms

Mistral Nemo 12B

12.2B

Q8

-

-

Mistral Nemo 12B

12.2B

FP16

-

-

Qwen 2.5 14B

14.7B

INT4

31 tok/s

~3436 ms

Qwen 2.5 14B

14.7B

Q8

-

-

Qwen 2.5 14B

14.7B

FP16

-

-

DeepSeek R1 32B

32.5B

INT4

-

-

DeepSeek R1 32B

32.5B

Q8

-

-

DeepSeek R1 32B

32.5B

FP16

-

-

Llama 3.3 70B

70.6B

INT4

-

-

Llama 3.3 70B

70.6B

Q8

-

-

Llama 3.3 70B

70.6B

FP16

-

-

Qwen 2.5 72B

72.7B

INT4

-

-

Qwen 2.5 72B

72.7B

Q8

-

-

Qwen 2.5 72B

72.7B

FP16

-

-

DeepSeek V3 671B

671B (37B active)

INT4

-

-

Need custom layer dimensions, offloading configurations, or distributed multi-node topology?

Open in Advanced VRAM Calculator

Inference Feasibility

Maximum model parameter class runnable by weight precision.

Context:

4-bit

Q4

32B Models

8-bit

Q8

32B Models

16-bit

FP16

14B Models

Fine-Tuning Feasibility

Hardware capacity for local training and adapter fine-tuning.

Context Length:

QLoRA

4-bit

32B Models

LoRA

16-bit

14B Models

Full Parameter Training

7B-8B Models

Workload Recommendations

Guidance for optimal precision and operational ceilings on NVIDIA RTX 6000 Ada Generation (48GB).

Inference Sweet Spot

Optimized for 14B models at uncompressed or Q8 precision, and 32B models at Q4 with up to 32k context.

Bandwidth & Speed Profile

With 960 GB/s aggregate bandwidth, batch size 1 inference operates in a memory-bandwidth bound regime. Generates approximately 213 tok/s on an 8B Q4 model and 25 tok/s on a 70B Q4 model.