ApX logoApX logo

LLM Inference: VRAM & Performance Calculator

Precision for model weights during inference. Lower uses less VRAM but may affect quality.

KV Cache precision. Lower values reduce VRAM, especially for long sequences.

▶ Runtime settings

Hardware Configuration

Select your GPU or set custom VRAM

Num GPUs

1

Devices for parallel inference

Input Parameters

Batch Size:

1

Inputs processed simultaneously per step

1
2
4
6
8

Sequence Length:

MHA
1,024

Context window size. Includes input and output tokens.

8K
16K
33K
66K
131K

Concurrent Users:

1

Number of users running inference simultaneously

1
2
4
6
8

▶ Advanced Configuration

Inference Simulation

(FP16 Weights / FP16 KV Cache) on 16GB / 360 GB/s / 26 TFLOPS Custom GPU

Input sequence length: 1,024 tokens

Configure model and hardware to enable simulation

Performance & VRAM Results

0.0%

VRAM

Ready

0 GB

of 12 GB VRAM

Generation Speed: ...

Time to First Token: ~0ms

Est. GPU Rental: N/A