Precision for model weights during inference. Lower uses less VRAM but may affect quality.
KV Cache precision. Lower values reduce VRAM, especially for long sequences.
▶ Runtime settings
Hardware Configuration
Select your GPU or set custom VRAM
Num GPUs
Devices for parallel inference
Input Parameters
Batch Size:
Inputs processed simultaneously per step
Sequence Length:
Context window size. Includes input and output tokens.
Concurrent Users:
Number of users running inference simultaneously
▶ Advanced Configuration
(FP16 Weights / FP16 KV Cache) on 16GB / 360 GB/s / 26 TFLOPS Custom GPU
Input sequence length: 1,024 tokens
Configure model and hardware to enable simulation©2025 ApX Machine Learning
0.0%
VRAM
0 GB
of 12 GB VRAM
Generation Speed: ...
Time to First Token: ~0ms
Est. GPU Rental: N/A
Assistant
Online